LatentEval
Eval rigor

Model routing and the refusal tax: a pre-registered study

On short, checkable tasks we measured no Opus 4.8-to-Fable 5 capability separation at either effort we ran, and the premium tier’s refusal ‘rescue’ silently served the cheaper model on 20 of 28 calls.

In brief

5 POINTS
  • A pre-registered benchmark on short, machine-checkable tasks found no measurable Opus 4.8 to Fable 5 capability separation at either effort tested, an inconclusive null.
  • Fable’s refusal “rescue” fired on 20 of 28 low-effort calls and every rescued answer was served by Opus 4.8, so the premium tier’s insurance mostly buys the cheaper model.
  • Refusals hit 75 to 86% of the benchmark’s benign tasks (21/28 low effort, 24/28 extended-high), against 16.25% on the wider v2 pool: the refusal tax tracks the task mix.
  • With rescue on, the pass rate reached 96.4% (27/28, 95% Wilson CI 82.3 to 99.4%), while Fable’s raw pass collapsed to 14 to 18% once refusals count as failures.
  • Route short, checkable work to the cheaper tiers, which lead on quality-per-dollar and reliability@k; reserve the premium tier for long-horizon work this study did not measure.
On this page (9)
Key figures (3)
Preprint · first release · Series 01 Pre-registered

Fable’s refusal rescue is sold as insurance. On this task family it mostly buys you the cheaper model at premium price: switch the server-side fallback on and most declined calls are silently answered by Opus 4.8. The Fable→Opus fallback fired on 20 of 28 low-effort calls, none stayed refused, and the pass rate with rescue on was 96.4% (27/28 trials, 20 of them rescued by Opus 4.8; 95% CI 82.3–99.4%). Behind that sits a benign-task refusal tax: on a task set that lands in Fable’s cyber-classifier failure zone, Fable declined roughly 75–86% of these machine-checkable tasks at both low and extended-high effort (21/28 vs 24/28; the standard ‘high’ tier was stop-ruled out, see Method), against 16.25% on the wider, lower-difficulty v2 pool. On capability we measured no separation between Opus 4.8 and Fable 5 at the two effort levels we tested (Δ −17.9pp at low, +3.6pp at extended-high; the pre-registered bar was unmet in both directions), an inconclusive null.

Bar: with Fable 5's fallback on, 20 of 28 declined calls were re-served by cheaper Opus 4.8, lifting pass rate to 96.4%. You pay premium for Opus.
”Rescuing” Fable’s refusals means serving the cheaper Opus 4.8 on 71% of callsServer-side fallback: on a policy refusal, Fable’s call is re-served by Opus 4.8 (Opus is ~half Fable’s token price).Source: data.json v3.withRescue (fallbackFiredCount 20, trials 28, rescuedPassRate 0.9643 [82-99], stillRefusedCount 0).

Headline findings

Rescue fired

20 / 28

Fable→Opus fallback fired on 20 of 28 low-effort calls; 0 stayed refused; pass rate with rescue on 96.4% (27/28, 95% CI 82.3–99.4%), 20 of 28 rescued by Opus 4.8.

Refusal tax

75–86%

Benign-task refusals at both low (21/28) and extended-high (24/28) effort: high at both effort levels; the low→xhigh difference is not distinguishable at this n. Wider v2 pool (40 tasks): 16.25% (13/80).

Capability separation

−17.9 / +3.6 pp

Fable minus Opus capability at low / xhigh. Pre-registered bar (≥15pp AND disjoint 95% CIs) unmet in both directions: inconclusive.

Fable raw pass

14–18%

User-experienced pass collapses to 17.9% (5/28) low and 14.3% (4/28) xhigh (95% CI 7.9–35.6% low) while capability on non-refused trials stays high; the gap is the tax.

Rescue is silently buying Opus

Turn on Fable’s refusal rescue and you mostly pay premium price for the cheaper model. Fable’s advertised protection against refusals is a server-side fallback: when a call is declined, the request is re-issued to Opus 4.8. With that fallback on at low effort, the fallback fired on 20 of 28 calls, zero calls stayed refused, and the pass rate with rescue on was 96.4% (27/28, 95% CI 82.3–99.4%), 20 of the 28 trials rescued by Opus. Every one of those 20 rescued answers was produced by Opus.

So the “insurance” you buy with the premium tier is, on this task family, largely just Opus. Fable’s published per-token price is roughly 2x Opus’s; when the rescue path carries most of the load, you pay that 2x for the cheaper model’s answer. The finding reconciles independently, too: the rescue arm logged 20 fallbacks against the plain sweep’s 21 refusals on the same 28 trials, consistent within run-to-run variation; the 20-versus-21 difference of one is itself the non-determinism.

A high refusal tax at both effort levels

Fable refused 75% of trials at low effort (21/28) and 86% at extended-high (24/28). Both effort arms are the same 14 tasks, so the comparison is paired, not independent, and a prompt that trips the classifier tends to trip it at both effort settings; the 75%→86% difference is well within run-to-run variation at this sample size, so we do not claim effort raises, lowers, or leaves the rate unchanged. What the runs show is the magnitude: the refusal rate is high at both effort settings, and at neither does Fable answer most of these benign tasks.

That number only means something when stated in the same breath as the task mix. This v3 set is a mix of reasoning, format, and tool-use problems with no facial cyber content: a seating-arrangement logic puzzle, a digit-sum-then-modular-exponentiation chain, a DP tiling count, word-capitalization format rules, and a wire-transfer tool-use approval. So 75–86% is specific to this failure-zone mix; the wider, lower-difficulty v2 pool refused just 16.25% (13/80). “Cyber” here is Anthropic’s own high-risk category label, applied broadly by the classifier (44 of Fable’s 45 v3 refusals carried it, as did all 13 of v2’s). It flags the request’s risk category and says nothing about the task content, which is why we read the tax as classifier over-triggering rather than a capability limit.

This elevated refusal rate is a known trade-off Anthropic disclosed after the Fable 5 redeploy: the retrained safety classifier flags more benign coding and debugging work. That makes it documented, expected behavior rather than a defect we uncovered, and the instrumentation for catching it is its own piece.

Bar chart: Fable 5 refused most benign checkable tasks, 75% at low and 85.7% at extended-high effort, while Opus 4.8 refused none. The gap is the tax.
The refusal tax: Fable’s raw pass rate collapses, but not because it failsEvery trial classified. Raw pass counts refusals as misses; capability excludes them (non-refused only).Source: data.json v2.{raw,capability,safetyRefusalRate} and v3.perModelPerEffort / refusalTaxByEffort.Capability on the non-refused sample; v3 Fable n=7 and n=4 are tiny (* = 95% Wilson CI shown). Deterministic exact-match / unit-test scoring.

Capability: an inconclusive null

On capability (the pass rate computed only over non-refused trials) we measured no separation between Opus 4.8 and Fable 5 at the two effort levels we tested. The pre-registered separation bar (a point gap of at least 15 percentage points AND disjoint 95% Wilson intervals) was unmet in both directions: Δ(Fable−Opus) was −17.9pp at low and +3.6pp at extended-high. Opus’s own rate moved from 89.3% (25/28) to 96.4% (27/28) across effort, a change within run-to-run variation on the same paired tasks. On the between-tier question, we could measure no separation on the n=4–7 non-refused samples. That outcome is inconclusive; it does not show the tiers are equal.

The word “capability” is load-bearing here. What a user actually experiences, raw pass with refusals counted as failures, is not tied: Fable’s raw pass collapses to 17.9% (5/28) at low and 14.3% (4/28) at extended-high, against Sonnet’s 96.4% (27/28) and Opus’s 89.3–96.4%. The gap between Fable’s high capability and its low raw pass is the refusal tax. And the capability number itself is fragile: Fable’s non-refused sample is only n=7 at low (71.4%, 95% CI 35.9–91.8%) and n=4 at extended-high (100%, 95% CI 51.0–100%). Those points are directional at best, never a “win” over Opus.

Capability pass rate on non-refused trials with 95% Wilson intervals: all overlap, so no measured Opus 4.8 vs Fable 5 separation on short tasks.
No daylight at the capability ceiling: all 95% intervals overlapCapability = pass rate on non-refused trials. Verdict: no measured Opus-Fable separation.Deterministic short-horizon tasks only. Long-horizon agentic & open-ended quality are NOT measured here.Source: data.json v3.perModelPerEffort.capabilityPassRate + capabilityPassRateCI95. Separation rule (≥15pp gap AND disjoint Wilson intervals) NOT met.* Fable points sit on n=7 (low) and n=4 (xhigh) non-refused: treat as indicative only.

Route checkable work to the cheaper tiers

On checkable work the cheaper tiers dominate on quality-per-dollar, and the premium’s price buys refusals rather than correctness. The figure below is the unitless quality-per-dollar ratio from the prior round (v2, n=40) on sticker pricing: Sonnet ~387, Opus 187, Fable 101; Sonnet leads Fable by roughly 3.8x.1 Fable’s published per-token price is about 2x Opus’s, so on this task family that 2x premium buys the refusal tax; capability does not rise with it.

Bar chart of quality per dollar, prior 40-task round: Sonnet 5 about 387, Opus 4.8 187, Fable 5 101; the cheapest tier buys the most per dollar.
The cheapest tier buys the most quality per dollar (deterministic short-horizon tasks)v2 40-task suite, 2 runs/model. Sticker pricing. Higher is better.Source: data.json v2.qualityPerDollar (opus 187.0, fable 100.7) = v2.raw ÷ v2.meanCostPerTaskUsd; Sonnet shown at sticker (~387).Fable’s lower value is the OPERATIONAL cyber-false-positive refusal tax (benign refusals scored as misses), not lower capability.Sonnet had INTRODUCTORY pricing ($2/$10 per Mtok, through 2026-08-31): quality/$ ~580, lead over Fable ~5.8×, vs ~3.8× at sticker $3/$15.

The reliability@k result points the same way: the premium tier did not buy run-to-run consistency. Over the 14 hardest items Opus scored 0.875 (low) and 0.946 (xhigh) and Sonnet 0.946 (low), against Fable’s raw 0.16, which is refusal-dominated and corroborates the refusal tax rather than standing as independent evidence; Fable’s capability@k of 0.75→1.0 sits only on the same non-refused sample (n=4 tasks at low, n=2 at extended-high). For short, machine-checkable work, route to the cheaper tiers, and reserve the premium tier for the long-horizon, open-ended work this study does not measure. The Sonnet routing call rests on lighter evidence: a low-effort spot-check plus the v2-round quality-per-dollar figure, and it was not part of the pre-registered Opus-vs-Fable design. The everyday version of this call, whether you personally need Fable 5, and the fuller Claude model choice both point the same direction.

Routing matrix: send deterministic checkable work to the cheaper tiers, which show no separation; reserve the premium tier for untested long work.
Routing by task type: what this suite measured, and what it did notColumn tint follows the model (Sonnet blue, Opus teal); the verdict is each cell’s label · hatched = outside this suite’s scope.† Rows 6–7 are NOT tested by this deterministic exact-match / unit-test suite; “premium may pay” is a hypothesis, not a result.Measured rows: data.json v2 / v3 (no measured Opus–Fable capability separation; Fable is refusal-prone on cyber-adjacent prompts).

Hypothesis ledger

Every hypothesis we pre-registered, and what the data returned. We predicted four separations; the data refused all four. Two rounds came back inconclusive.

TABLEShow full table (6 rows)Showing full table (6 rows)
HypothesisWhat we expectedWhat we foundVerdict
v1 (origin): the cheapest tier is capability-limited on this task pool.Cheaper models fail measurably more than the premium tier.Ceiling effect: the cheapest tier scored at or near 100% on the hardest items (an unsized ceiling observation); the pool was too easy to discriminate. inconclusive
v2 (bridge): harder, research-informed traps separate the tiers.The premium tier pulls ahead on an n=40 harder set.Three tiers capability-tied (Sonnet 92.5%, Opus 98.75%, Fable 98.51%; overlapping Wilson CIs). The one clear effect was Fable’s 16.25% benign-task refusal tax. not-supported
H1 (v3, pre-registered): Fable’s capability minus Opus’s increases with effort.The premium’s headroom appears once you let it think.Δ(Fable−Opus) = −17.9pp (low), +3.6pp (xhigh); the intermediate high arm was not run per the stop rule. Opus’s own 89.3%→96.4% move is within run-to-run variation on the paired sample. not-supported
H2 (v3, pre-registered): on items where Opus fails at low effort, Fable passes materially more often at high/xhigh.Fable clears the Opus-failure-zone once effort is high.No separation; Fable’s capability rides a tiny non-refused sample (5/7 low, 4/4 xhigh) with very wide CIs. not-supported
H3 (v3, pre-registered): where pass-rate ties, Fable’s reliability@k on the hardest items exceeds Opus’s (the premium buys consistency).Fable is more run-to-run consistent on the hardest items.Opus 0.875–0.946 ≥ Fable’s raw reliability@k of 0.16; Fable’s capability@k (reliability@k over non-refused trials only) of 0.75→1.0 sits on that tiny sample. not-supported
Core test: an Opus–Fable capability separation exists at some effort (the study’s central question).A detectable ≥15pp capability gap with disjoint 95% CIs at some effort.No separation in either direction; Fable’s capability CI is enormous on the n=4–7 non-refused sample. Raw pass is not tied: Fable’s collapses to 14–18% via refusals. inconclusive

Method

The study is deliberately narrow: 14 pruned Opus-failure-zone tasks, two runs each, 28 trials per arm at two effort points. It ran in three rounds against a single provider’s tiers (Anthropic) on the Anthropic API. Round v1 (origin) was inconclusive by a ceiling effect (the tasks were too easy to discriminate tiers), which motivated harder sets. Round v2 (bridge) ran n=40 tasks at 80 trials per model. Round v3 is the discriminating sweep those rounds motivated, run at low and extended-high effort; the intermediate “high” arm was not run, because a pre-registered stop rule fired once no Opus–Fable gap emerged at low or extended-high, so we did not spend runs chasing one. The refusal “rescue” arm (Fable with the server-side Fable→Opus fallback enabled) was part of the pre-registered design and ran at low effort; it is what the hero reports.

Scoring is deterministic: exact-match, unit-test, and tool-argument matching, with no LLM judge. Every rate carries a 95% Wilson interval, and tier comparisons use two-proportion tests. Pricing figures use sticker pricing as the primary basis, from a 2026-07-02 snapshot of public per-token prices.

This study’s git-verified pre-registration means the design, hypotheses, and stop rule were committed before any data existed, with the Outcome field left as TBD at design time and filled only after the run. It is an internal, auditable record checkable against the repository history; it was never lodged with an external registry. Every count, rate, and confidence interval on this page recomputes from the results CSV; the quality-per-dollar ratios derive from those counts and public per-token pricing.

Metrics and definitions

We separate metrics we borrowed from standard practice from the ones we operationalized ourselves.

  • Capability pass rate: borrowed from pass@1 / exact-match accuracy on non-refused trials. Ours: the fraction of non-refused trials scored correct by deterministic exact-match / unit-test / tool-argument matching (no LLM judge); refusals are excluded as an operational rather than capability outcome; denominator = trials − refusals; 95% Wilson intervals throughout.
  • Raw pass rate: borrowed from task success as a user experiences it. Ours: the fraction of all trials correct, with refusals counted as failures (denominator = 28). The gap from capability is the refusal tax.
  • reliability@k: borrowed from the pass@k / pass^k consistency lineage (e.g. τ-bench pass^k). Ours, with the formula published so a reader can reproduce it: the mean over tasks of (per-task pass fraction)^k, i.e. run-to-run consistency, distinct from the naive all-runs-pass rate. Where we write capability@k, it is the same measure computed over non-refused trials only (consistency with refusals set aside).
  • Refusal tax: our label for a safety-refusal rate the API returns (an HTTP 200 with empty content on a benign task). We report it as a rate at each effort level, always with the task mix stated alongside, and because the two effort arms are the same tasks we do not read the low-to-xhigh difference as an effort trend in either direction.
  • Separation criterion: borrowed from two-proportion significance testing plus a minimum effect-size threshold. Ours: Opus and Fable are declared separated on a dimension at an effort level iff the point difference is at least 15 percentage points AND the two 95% Wilson intervals are disjoint. Fixed before any data (pre-registered).
  • Quality-per-dollar: borrowed from cost-efficiency ratios. Ours: a unitless efficiency ratio from the prior round (v2, n=40), reported on sticker pricing as primary with the introductory-pricing basis footnoted, and kept out of the machine-readable CSV as is any raw dollar figure.

Scope and limitations

  • Small n on the discriminating arm. The v3 effort sweep is 14 pruned Opus-failure-zone tasks × 2 runs = 28 trials per arm (the prior round was n=40). Fable’s capability points ride n=5/7 (low) and 4/4 (xhigh) non-refused trials with very wide Wilson intervals, directional only and unable to rank Fable against Opus.
  • Task-mix-driven refusal rate. The 75–86% refusal rate is specific to this failure-zone set, a mix of reasoning, format, and tool-use problems with no facial cyber content (a seating-arrangement logic puzzle, a digit-sum-then-modular-exponentiation chain, a wire-transfer tool-use approval). The wider, lower-difficulty v2 pool was 16.25% (13/80). Always read the number with the mix.
  • Two effort points only. Low and extended-high; the intermediate “high” arm was not run per the pre-registered stop rule, and thinking is non-deterministic. Refusals are high at both effort levels (21/28 low, 24/28 xhigh); because the two arms are the same tasks, the difference is within run-to-run variation and no effort trend is claimable in either direction.
  • “Cyber” is the API’s label; “false positive” is our reading. Both rounds carry Anthropic’s own per-refusal policy category: all 13 of v2’s refusals, and 44 of v3’s 45 sweep refusals, came back categorised “cyber.” Calling them false positives is our judgment that the underlying tasks (reasoning, format, and tool-use problems with no facial cyber content) are benign; a reader can verify the task set.
  • Calibration truncations not re-validated. The 14-task Opus-failure-zone set was selected from a calibration pass that logged roughly 12% max_tokens truncations, and that selection counted truncated failures the same as reasoning failures. Calibration was not re-run after the per-category caps were corrected (the sweep itself had zero truncations), so the task set’s stability under the corrected caps is not re-validated here.
  • Pre-registration is internal. It is git-verified against repository history (design + hypotheses + stop rule precede the data/outcome commit, with Outcome=TBD at design time), an internal record rather than an external lodgement.
  • Pricing is a moving target. Primary figures use sticker pricing; Sonnet’s quality-per-dollar lead uses introductory pricing (valid through 2026-08-31) and narrows from about 5.8x to about 3.8x on sticker. v3’s per-task quality-per-dollar inverts for measurement reasons and is not shown; the footnote explains why.
  • A single snapshot. One provider, one task family, one 2026-07-02 pricing and classifier snapshot; classifier behavior and prices can shift and move every number here.

This site is building toward a reliability profiler meant to run sweeps like this one and return the routing verdict with its confidence interval attached. The instrument has not launched, so it contributes no figure here; every count, rate, and interval above recomputes from the results CSV.

Footnotes

  1. On Sonnet’s introductory pricing (valid through 2026-08-31) the quality-per-dollar figure is about 580 and the lead over Fable about 5.8x; sticker pricing (used here as the primary basis) narrows it to about 3.8x. v3’s own per-task quality-per-dollar inverts (Fable appears cheaper) for two reasons that both cut against a value reading: refusals return short outputs that deflate Fable’s per-task cost, and the numerator definition itself changed between rounds (v2 scored raw-pass rate, v3 scored capability pass rate), so the v3 ratio is not comparable to v2’s. It is a deflation artifact that carries no value signal, so it is not shown.

Sources

  1. Model routing and the refusal tax: results dataset (CSV) Retrieved
  2. Edwin B. Wilson (1927), Probable Inference, the Law of Succession, and Statistical Inference (JASA) Published
  3. τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains Published
  4. Evaluating Large Language Models Trained on Code (Codex) Published
  5. Anthropic: Refusals and fallback Retrieved
  6. Anthropic: Claude pricing Retrieved

How to cite

LatentEval. "Model routing and the refusal tax: a pre-registered study." 2026. https://latenteval.ai/research/model-routing-refusal-tax

@misc{model-routing-refusal-tax-2026,
  author = {{LatentEval}},
  title = {Model routing and the refusal tax: a pre-registered study},
  year = {2026},
  url = {https://latenteval.ai/research/model-routing-refusal-tax}
}