Skip to content
All posts
· 8 min read · Damiën Semler

A hidden tool can still be called

An MCP study shows why public tool metadata, caller-specific listings and direct execution checks need separate evidence.

A tool can disappear from an agent's menu and remain callable at the protocol endpoint. That distinction is the central result of a September study of MCP authorization: filtering tools/list changed what model-driven clients attempted, while a scripted client could still invoke a hidden tool when the server lacked a separate execution check.

For agent registries, this creates a precise measurement problem. A catalog can describe a capability and its advertised permission. It cannot infer from a listing that the permission is enforced on a direct call. Discovery, visibility, and execution are three different observations, each with a different evidentiary limit.

A catalog is part of the runtime

The MCP 2026-07-28 release made list responses cacheable and put method and tool names into HTTP headers for routing and metering. That makes a tool catalog operational infrastructure: a client may reuse a list rather than fetch it for every call, and an intermediary can identify the requested method before parsing the body. Neither change makes the list an authorization decision.

MCP Apps' authorization guidance documents two patterns: protecting an entire server or requiring authorization for particular tools. In the latter pattern, public calls can proceed without a token, while a protected call receives an HTTP 401 challenge and is retried after authorization. The protected tool handler should also check its authentication context. The guidance's separation of a visible tool from its execution gate is important: a client can learn a tool's name through documentation, a cached list, a log, or another party.

A recent enterprise MCP paper by Huan Li, Yuwei Wang and Srinivasan Manoharan examines what happens when those surfaces use different permission rules. The authors propose deriving anonymous registry metadata, each authenticated caller's component listing, and the invocation check from one server-side declaration. This is an implementation design, not a new MCP requirement.

What the experiment actually shows

The researchers used an eight-tool mock Kubernetes server, with four viewer tools and four admin tools. They compared three configurations: all tools listed with a permission check inside the admin tool body; admin tools hidden from viewers but without an execution check; and admin tools hidden with an execution check. Across four models and 720 prompts per configuration, the first setup elicited 152 attempts to call the forbidden restart_pod tool. The filtered setups elicited zero such model-issued tool calls. Li et al., evaluation and Table 4.

That zero has a narrow interpretation. The model could not emit a call to a tool absent from the schema it had been shown. In the baseline, the 152 attempts were denied by the tool-body check; they were not 152 successful intrusions. The paper does not establish that one model is generally safer than another, or that a larger attack corpus would preserve these exact attempt rates. Its 60 payloads and tool implementation were a controlled testbed. Li et al., methods and limitations.

The second test exposes the missing control. A scripted client sent a direct tools/call using the hidden tool's name. The visibility-only configuration executed it. Both configurations with an invocation check rejected it. Hiding a tool changed the agent's menu; it did not revoke the endpoint's capability. Li et al., Table 6.

Three surfaces of a tool permission An anonymous registry reads descriptive metadata. An authenticated agent receives a permission-filtered tool list. A direct call must be checked again before execution. Description and visibility do not grant execution. One permission rule, three different observations The server must enforce the last step, even when earlier steps are skipped. Registry metadata What exists? What is declared? Descriptive evidence Caller listing What may this caller see? Which schema is shown? Visibility evidence Direct invocation May this caller run it? Is the call denied? Enforcement evidence A hidden tool may still be named in a raw tools/call request.

The registry's least-privilege tradeoff

The paper's proposal includes a public metadata endpoint for registries. It describes tool schemas and required permissions without granting a registry credentials to invoke those tools. This avoids centralizing execution credentials across every indexed server. The cost is also explicit: public metadata can reveal tool names and permission strings that help an attacker plan a direct call. The researchers did not measure whether that reconnaissance increases prompt-injection success; they identify it as a reasoned risk. Li et al., Section 4.3.5.

A registry therefore needs careful labels. “Advertised” can mean the server declared a tool or a required scope. “Observed for caller” can mean an authenticated listing included or omitted it under a particular credential. “Enforced” requires an authorized test of what happens when that caller invokes it directly. An anonymous crawler generally cannot perform the latter for a protected tool. It should publish unknown, not silently promote a declaration into a security verdict.

There is a useful parallel in A2A v1.0.1. An Agent Card advertises skills and authentication schemes, while the specification assigns servers responsibility for authenticating requests and scoping access to tasks and resources. A public card is evidence of a declaration; it cannot by itself show whether a private task is isolated from another caller. The protocols differ in their objects and operations, but the measurement boundary is the same: published capability data and request-time authorization answer different questions.

A small, testable contract

For operators, the paper suggests a concrete contract: generate public metadata, caller-specific listings, and call-time decisions from one permission source. Then test all three against the same credential matrix. At minimum, include an anonymous caller, a caller with a narrow scope, a caller with broader scope, and a direct call to a tool hidden from the narrow caller. The direct call is the decisive negative test.

For registries and verifiers, the corresponding contract is about wording. Attach a principal, a timestamp, and an observation method to every visibility claim. Keep operator-supplied permission metadata separate from observed behavior. Record when a protected test was impossible. An impressive catalog coverage figure says little about execution controls unless the test crossed that boundary with permission.

What we are watching

The open implementation question is how often production servers keep discovery and enforcement aligned as tools, roles, and caches change. The new MCP cache hints make stale visibility worth testing explicitly; a tool removed from a fresh listing may still appear in a client's cached schema until that cache expires. The security decision still belongs at invocation.

A useful next measurement would report three separate rates: servers that declare permissions, servers that vary listings by caller, and servers that reject unauthorized direct calls. Conflating them would make the easiest surface to crawl stand in for the hardest one to verify.

Sources

Cite this post

@misc{semler2026hiddentool,
  author = {Semler, Damiën},
  title = {A hidden tool can still be called},
  year = {2026},
  month = oct,
  url = {https://agenstry.com/blog/hidden-tool-still-callable},
  publisher = {Agenstry}
}
← Back to blog Agenstry