LatentEval

INSTRUMENT | eval

MCP Tool Surface Auditor

1 cited source

Audit an MCP tool surface or diff two manifests for breaking changes. Paste your tools/list output. Nothing from the manifest you paste is sent anywhere.

Paste your tools/list output and see how large your tool surface is, which tools the model is likely to confuse, and which parameters it has to guess at. The MCP gateway guide explains what consolidation buys and what it risks; this page reads a surface the guide cannot. For the token budget those schemas consume inside a context window, see the token counter.

What are you doing

JSON. Drag one here, or use the box below.

Accepts a JSON-RPC response, a bare {tools: [...]} object, a plain array of tool definitions, a map of server names to any of those, or OpenAI and Anthropic format tool arrays. Nothing from the manifest you paste is sent anywhere.

Check this value.

J(A, B) = |A ∩ B| / |A ∪ B|How?

How this is calculated

The size column measures the whole tool definition, including name, title, description and schema, not the input schema alone. Every size is taken on one canonical serialization, so the same tool delivered through any accepted shape measures the same.

Accepted shapes: a JSON-RPC tools/list response, a bare {tools: [...]} object, a plain array of tool definitions, a map of server names to any of those, and vendor tool arrays. The vendor formats named here are OpenAI-style function and tool definitions, whose schema key is parameters, and Anthropic-style tool definitions, whose schema key is input_schema. Both normalize to inputSchema before anything is measured.

Overlap scoring. Names are split at each camelCase boundary and at every run of non-alphanumeric characters, then lowercased. Descriptions are lowercased, split the same way, and reduced by dropping single-character tokens and this stopword list of 24 words: a an the in of for to and or on by with from at is be this that it as are its your you. A pair is flagged when the name score reaches 0.50 or the description score reaches 0.40. Both thresholds are judgment calls with no primary source, which is why they are printed here beside the list they run on. When both token sets are empty the score is 0, never 1 and never NaN. Above 1,000 tools the pair comparison is skipped; counts and sizes are unaffected.

Diff classification. The organizing principle is the direction of change: narrowing what a caller may send is breaking, because a call that was valid before can now be rejected; widening it is not, because everything the caller already sends stays valid. Anything this version does not rank is still reported, as a warning saying it was not classified, because a change shown as nothing is the worse failure.

Output schemas are the one place this version is deliberately conservative. For responses the direction inverts: a producer widening what it may return breaks consumers who narrowed their expectations. Version 1 classifies every outputSchema change as a warning and never as breaking. The sharp case, a name removed from outputSchema.required or a property removed from outputSchema.properties, keeps that class and is collected separately as a withdrawn output guarantee.

Token counts are exact or absent. They come from the o200k_base byte-pair table under Sources, loaded only when you ask for them; there is no heuristic estimate, because JSON with braces, quotes and snake_case identifiers tokenizes differently from the prose those estimates are calibrated on. Retrieval over a large tool surface has been measured to matter: Gan and Sun report tool-selection accuracy of 43.13 percent against 13.62 percent, cited in full below.

Where a surface carries tools that write, delete or spend, the irreversible action inventory is the page that sorts them by what cannot be undone. This one reads shape and size, not consequence.

Formula: J(A, B) = |A ∩ B| / |A ∪ B|

Questions

Why does the page refuse my config file?

A client config carrying command or env is not a tool surface, and saying so is better product than parsing it. The check reads key names only, never values, and stops before anything else is tried. This is what protects secrets in MCP client configs from being read at all.

What does 'lower bound' mean on the tool count?

A tools/list response that carries a nextCursor is one page of a longer list. This page reads a file and makes no network call, so it cannot fetch the next page. When a cursor is present every count is labeled a lower bound.

Why are renames listed as a removal plus an addition?

Two manifests carry no identity beyond the name. A rename is undecidable from them, so it is classified as a removal plus an addition. Where a removed tool and an added tool score at or above 0.50 on names and carry a byte-identical input schema, the page offers a labeled hint. The hint is never a verdict and never changes the breaking count.

What is a silent redefinition?

A tool whose description changed while its schema stayed byte-identical. Nothing about the call signature moved, so no compatibility check fires, but the model is now being told the tool does something else. The page reports it as a warning in its own class.

Why is the reported size the whole definition and not just the schema?

The size column measures the whole tool definition, including name, title, description and schema, not the input schema alone. What the surface pays for is the tool definition's footprint, not the input schema alone, so that is what is measured. If you are comparing this figure against your own count of inputSchema bytes, that is the difference.

Sources

  1. RAG-MCP: Mitigating Prompt Bloat in LLM Tool Selection via Retrieval-Augmented Generation: the measured tool-selection accuracy of 43.13 percent against 13.62 percent, the one external claim this page carriesarXiv Retrieved