INSTRUMENT | eval
MCP Tool Surface Auditor
1 cited source
Audit an MCP tool surface or diff two manifests for breaking changes. Paste your tools/list output. Nothing from the manifest you paste is sent anywhere.
Paste your tools/list output and see how large your tool surface is, which tools
the model is likely to confuse, and which parameters it has to guess at. The
MCP gateway guide explains what consolidation buys and what it
risks; this page reads a surface the guide cannot. For the token budget those schemas consume
inside a context window, see the
token counter.
Tools
Breaking changes
Surface summary
| Measure | Value |
|---|
Per-tool sizes
Definition size (chars / bytes / tokens)
| Tool | Server | Chars | Bytes | Tokens |
|---|
Overlap
| Tool A | Tool B | Name J | Desc J | Shared tokens | Fired on |
|---|
Missing descriptions
| Measure | Value |
|---|
Breaking changes
| Tool | Field | What changed |
|---|
Warnings
| Tool | Field | What changed |
|---|
Compatible changes
| Tool | Field | What changed |
|---|
J(A, B) = |A ∩ B| / |A ∪ B|How?
How this is calculated
The size column measures the whole tool definition, including name, title, description and schema, not the input schema alone. Every size is taken on one canonical serialization, so the same tool delivered through any accepted shape measures the same.
Accepted shapes: a JSON-RPC tools/list response, a bare {tools: [...]}
object, a plain array of tool definitions, a map of server names to any of those, and vendor
tool arrays. The vendor formats named here are OpenAI-style function and tool definitions,
whose schema key is parameters, and Anthropic-style tool definitions, whose schema
key is input_schema. Both normalize to inputSchema before anything is
measured.
Overlap scoring. Names are split at each camelCase boundary and at every run of non-alphanumeric characters, then lowercased. Descriptions are lowercased, split the same way, and reduced by dropping single-character tokens and this stopword list of 24 words: a an the in of for to and or on by with from at is be this that it as are its your you. A pair is flagged when the name score reaches 0.50 or the description score reaches 0.40. Both thresholds are judgment calls with no primary source, which is why they are printed here beside the list they run on. When both token sets are empty the score is 0, never 1 and never NaN. Above 1,000 tools the pair comparison is skipped; counts and sizes are unaffected.
Diff classification. The organizing principle is the direction of change: narrowing what a caller may send is breaking, because a call that was valid before can now be rejected; widening it is not, because everything the caller already sends stays valid. Anything this version does not rank is still reported, as a warning saying it was not classified, because a change shown as nothing is the worse failure.
Output schemas are the one place this version is deliberately conservative. For responses the
direction inverts: a producer widening what it may return breaks consumers who narrowed their
expectations. Version 1 classifies every outputSchema change as a warning and never
as breaking. The sharp case, a name removed from outputSchema.required or a
property removed from outputSchema.properties, keeps that class and is collected
separately as a withdrawn output guarantee.
Token counts are exact or absent. They come from the o200k_base byte-pair table under Sources, loaded only when you ask for them; there is no heuristic estimate, because JSON with braces, quotes and snake_case identifiers tokenizes differently from the prose those estimates are calibrated on. Retrieval over a large tool surface has been measured to matter: Gan and Sun report tool-selection accuracy of 43.13 percent against 13.62 percent, cited in full below.
Where a surface carries tools that write, delete or spend, the irreversible action inventory is the page that sorts them by what cannot be undone. This one reads shape and size, not consequence.
Formula: J(A, B) = |A ∩ B| / |A ∪ B|
Questions
Why does the page refuse my config file?
A client config carrying command or env is not a tool surface, and
saying so is better product than parsing it. The check reads key names only, never values, and
stops before anything else is tried. This is what protects secrets in MCP client configs from
being read at all.
What does 'lower bound' mean on the tool count?
A tools/list response that carries a nextCursor is one page of a
longer list. This page reads a file and makes no network call, so it cannot fetch the next
page. When a cursor is present every count is labeled a lower bound.
Why are renames listed as a removal plus an addition?
Two manifests carry no identity beyond the name. A rename is undecidable from them, so it is classified as a removal plus an addition. Where a removed tool and an added tool score at or above 0.50 on names and carry a byte-identical input schema, the page offers a labeled hint. The hint is never a verdict and never changes the breaking count.
What is a silent redefinition?
A tool whose description changed while its schema stayed byte-identical. Nothing about the call signature moved, so no compatibility check fires, but the model is now being told the tool does something else. The page reports it as a warning in its own class.
Why is the reported size the whole definition and not just the schema?
The size column measures the whole tool definition, including name, title, description and schema, not the input schema alone. What the surface pays for is the tool definition's footprint, not
the input schema alone, so that is what is measured. If you are comparing this figure against
your own count of inputSchema bytes, that is the difference.