LatentEval

Glossary

Capability tier (model routing)

Capability tier is the band a router sorts a model into, ordered by how much task competence its vendor claims it delivers. The ordering is published as a product hierarchy, so whether a given boundary changes your results is a question only a paired eval on your own tasks can settle.

Capability tier is the band a model router sorts a model into, ordered by how much task competence its vendor claims it delivers. The bands run in a familiar order: a premium band for long, open-ended work, a middle band for general production traffic, a small fast band for classification and extraction. A router is the component that applies the ordering to each request; the tier is the ordering itself, published as a product hierarchy and priced accordingly.

Because a tier is an ordering rather than a measurement, the only thing you can report about it is what the boundary does to one specific task family. That report is a gap in pass rate between two adjacent bands, scored on the same items, with an interval on the difference and the denominator that produced it. A gap quoted without both of those restates the vendor’s ordering in your own numbers.

Our pre-registered routing study on short, machine-checkable tasks is the run this page is built on, and it found the boundary harder to see than the price list implies. Against a bar fixed in advance at a point gap of at least 15pp together with disjoint 95% Wilson intervals, it measured no capability separation between Claude Opus 4.8 and Claude Fable 5 at either effort level tested. That outcome is inconclusive rather than a finding of equality: the non-refused samples ran to n=4 to 7, so the intervals around them are very wide.

Treat a published tier as a hypothesis about your workload, and test it on paired items before you route by it.

How to test whether a tier boundary is real

Run both models over the same items, score each pass or fail, and read the paired table instead of two independent rates. Pairing removes the item-difficulty variance that otherwise swamps a between-model gap, so the same number of runs buys a much tighter interval on the difference. The McNemar test calculator takes the four paired cells and returns an exact p-value alongside a confidence interval on the pass-rate difference. Fix the separation bar before the run, stated as a point gap and an interval condition together, since a bar chosen after the data lands can be moved to wherever the data went. Size the run first with the sample size and power calculator, because a tier gap worth routing on is often small, and an underpowered run returns an inconclusive null that reads on the page like equality.

Report the denominator beside every rate you publish. Refusals, timeouts and truncations each remove a request before it can be scored. A band’s pass rate computed over the answers that came back is therefore conditioned on which requests that band agreed to answer, which is coverage conditioning operating on your comparison, so put answer coverage next to the rate. Instrument the measurement at the router, where a request still carries both the tier it was assigned and the model that actually served it, since a fallback chain can hand the call to a different band without changing the response status.

Capability tier vs usage tier

A usage tier is a billing and throughput band on an API account. Anthropic’s standard bands are named Start, Build, Scale and Custom, with an Evaluation tier below them that new organizations and organizations with little usage history may start in, and placement happens automatically from usage history and account standing. The first three bands each set a monthly spend cap, the Custom tier sets none and arranges its limits with an account team, and every band sets per-minute request and token ceilings.1 Capability tier is a claim about which model does better work. The two are independent: every model in the family is callable at every usage tier, and the rate limits are applied separately per model, so growing an account raises throughput without unlocking a better model.

The two orderings also run partly opposite each other. At the Start tier, Claude Fable 5 carries a 500,000 input-token-per-minute ceiling against the 2,000,000 allowed for Claude Opus 5, Sonnet 5 and Haiku 4.5 alike.1 The model a vendor prices at the top of its range can be the one with the least throughput to give you. A router that promotes bulk traffic into the premium band can collect a 429 the cheaper band would have absorbed. So the capability ordering tells you what a request deserves, while the usage tier sets what the account can actually push through in a minute.

Price stops predicting capability the moment the task is checkable

A price tier is the per-token rate card, the input and output prices that sort the same models into an ordering you can read off an invoice. Vendors publish the two orderings side by side, so price is usually taken as a legible proxy for capability. The proxy holds up on open-ended work and comes apart on checkable work, where a cheap model either produces the right answer or does not.

Our routing study is a worked case of it coming apart. On short, deterministic, machine-checkable tasks the premium band’s extra cost bought no measured capability over the cheaper one, while its refusal behavior dragged the delivered result down. Raw pass rate, counting refusals as the failures a user experiences, fell to 17.9% (5 of 28, 95% CI 7.9–35.6%) at low effort and 14.3% (4 of 28) at extended-high effort, against 89.3 to 96.4% for Opus 4.8 across the same paired items. The capability ordering and the delivered-outcome ordering separated, and refusal rate is what pulled them apart. The price tier tells you what a route costs, and only a paired run on your own items tells you what it returns. The task-by-task form of this decision sits in the guide to picking a Claude model, and what reaching up actually costs carries the evidence behind it.

Capability tier vs capability threshold in AI safety

A capability threshold in frontier-safety policy is a level of dangerous capability that, once a developer judges a model to have reached it, obliges that developer to apply a heavier set of safeguards. Anthropic’s Responsible Scaling Policy uses exactly that phrasing and pairs it with AI Safety Levels. Reaching a Capability Threshold requires upgrading to the ASL-3 Security Standard, with the thresholds themselves drawn in areas such as chemical and biological weapons uplift and automated AI research.2 That usage is established and precise, and it carries a meaning separate from the routing one.

The two kinds of band answer different questions. A routing tier ranks models by expected usefulness so a request can go to the cheapest one that will do the job, and crossing a boundary changes what you pay. A safety threshold ranks capability by what it would let a misuser accomplish, and crossing it tightens controls on training and deployment. A model can sit at the top of a vendor’s routing hierarchy without having reached any published safety threshold, and the two labels move on separate clocks.

No single capability number fixes a tier boundary for everyone, because the size of the gap between two bands is a property of the task family you run it on. The same boundary can be worth paying for on long-horizon agent runs, where an early error compounds through every later step, and vanish on short calls a script can check. Capability tier gives you the hypothesis; the measured gap on your own items, with its interval, tells you where the boundary falls. The instrumentation for catching a route that silently served a band other than the one you asked for is in logging refusals and model fallbacks, and the two-band version of the decision, rate cards included, is worked through in Opus against Sonnet. For a team routing one short task family whose answers a script can check, the routing study says where to look first: the premium band lost there on delivered outcome, on that particular task mix and against that vendor’s refusal classifier. Both of those conditions are narrow enough that neither travels on its own, so run the comparison on your own items and count refusals as failures while you do.

Footnotes

  1. Anthropic, “Rate limits,” Claude API documentation. States that limits are defined by usage tier (Start, Build, Scale, Custom, plus an Evaluation tier that new organizations and organizations with limited usage history may start in) with organizations placed automatically on the basis of usage history and account standing; that the Start, Build and Scale tiers each carry a monthly spend cap while organizations on the Custom tier have none; that rate limits are applied separately for each model; and lists the Start-tier ceiling for Claude Fable 5 at 500,000 input tokens per minute against 2,000,000 for Claude Opus 5, Sonnet 5 and Haiku 4.5: https://platform.claude.com/docs/en/api/rate-limits (as of 2026-08). 2

  2. Anthropic, “Anthropic’s Responsible Scaling Policy.” Defines Capability Thresholds whose crossing obliges an upgrade of safeguards to the ASL-3 Security Standard, with thresholds drawn in areas including chemical and biological weapons uplift and automated AI research and development: https://www.anthropic.com/responsible-scaling-policy (as of 2026-08).