LatentEval

INSTRUMENT | Eval statistics

Eval Budget Calculator: What an LLM Evaluation Run Will Cost

6 cited sources

Price an eval plan before you run it: cases by variants by trials by judge calls, plus the human review hours it needs, from your own rates.

An eval plan is a multiplication before it is a budget. This prices one from your own rates, on the evidence in our work on how many runs an eval needs and the statistics behind them.

Size it first: cases is the distinct-item count, trials the repeats per case, variants the number of arms. The sample size and power calculator sizes the set. Human review time is usually the largest line.

The plan

Distinct items in the set.

Check this value.

Arms under test: prompts, models, configurations.

Check this value.

Repeats of each case, per variant.

Check this value.

0 for a deterministic-only eval with no judge.

Check this value.

Machine cost
Price the judge

USD per judge call. What one judge call costs you, all in.

Check this value.

USD per judge call. Optional. Some eval platforms bill per scored item.

Check this value.

USD per graded item. Optional. What the system under test costs per output, before any judging.

Check this value.

Human cost

Percent of graded items a person reads.

Check this value.

How long one item takes to judge by hand.

Check this value.

USD per hour. Salary plus everything on top of it. No preset here: a loaded rate is local to your team.

Check this value.

Graded items

3,000

3,000 graded items and 6,000 judge calls at these counts. Enter your rates to price this plan.

Cost per case

not priced

Cost per graded item

not priced

These are your own rates. We filled nothing in.

Enter your rates to price this plan.

LineAmount
Graded items 3,000
Judge calls 6,000
Judge spend not priced
Run spend not priced
Machine spend not priced
Human review hours not priced
Human spend not priced
Total spend not priced

What the same plan costs at other trial counts

TrialsGraded itemsMachineHumanTotal spend
1 600 not priced not priced not priced
2 1,200 not priced not priced not priced
3 1,800 not priced not priced not priced
5 3,000 not priced not priced not priced
10 6,000 not priced not priced not priced
Export

total = judge calls · price + items · run cost + hours · rateHow?

How this is calculated

The multiplication. Cases times variants times trials gives the graded items, and each graded item takes however many judge calls you set. Nothing here is a fixed cost, so the total is exactly linear in all three counts: double the trials and you double the bill. That is the property the sensitivity table below the result reads off, and it is why deciding the counts is the budget decision. Sizing them is a question of statistical power, not of taste, and it belongs upstream of this page.

The worked example, small enough to check by hand. 200 cases, 3 variants and 5 trials is 3,000 graded items and 6,000 judge calls. At $0.004 a call plus a $0.0025 platform fee, judging costs $39.00; at $0.02 to produce each output, the run costs $60.00. Reviewing 10 percent of 3,000 items at 3 minutes each is 15.0 hours, which at $95 an hour is $1,425.00. Total $1,524.00, of which the human column is 93.5 percent. That last number is the one most plans get wrong.

Two ways to price the judge. Per call is what you use when you already know the blended cost of one judgment. Per token is what you use when you know the prompt: input tokens times the input price plus output tokens times the output price, divided by a million. Some vendors re-price a long request entirely, so a call above their threshold costs more per token across the whole call and not just above the line. Where a preset carries that rule, the page applies it and the stamp says so.

The rates are yours. Every money field starts empty. The presets are prices a person read from the vendor's own page on the date shown beside each one, offered because typing four of them is tedious, and they can be out of date the day after they were read. Your own number always wins, and editing a filled field turns the stamp back to your number.

Rounding. Everything computes at full precision and rounds only for display. Money shows two decimals, or four when a per-item figure is under a dollar, because a real sub-cent cost printed as $0.00 is worse than useless. Review time is never rounded on its way to hours: 10 percent of 3,005 items is 300.5 items and 15.025 hours. Full precision has one outer limit worth stating: totals are exact to the cent up to about $90 trillion, which is where double-precision arithmetic runs out of cents. The maxima on this form can be combined into a plan past that line, and the cents in such a total are approximate. No budget anyone brings here is within four orders of magnitude of it.

Honest limits. This is arithmetic on numbers you supply. It has no opinion on whether your plan is big enough to answer your question, and it does not know what your judge agrees with. How much human review to buy has a principled answer, sized against a power target rather than a budget, and it is a different calculation from this one. Treat the total as a floor: the lines under "What this does not count" below are all real money and none of them are in it.

Formula: total = judge calls · price + items · run cost + hours · rate

The dated rates

Per 1M tokens in US dollars, by the vendor's own model id. Above the threshold in the long-context column, the whole call bills at 2x input and 1.5x output. Gemini 3.7 Flash is due for a re-read by 2026-12-31

ModelInputOutputLong-context ruleWe read it
claude-sonnet-5 $2.00 $10.00 none 2026-08-24
claude-haiku-4-5 $1.00 $5.00 none 2026-08-24
claude-opus-5 $5.00 $25.00 none 2026-08-24
gpt-5.6-terra $2.00 $12.00 above 272,000 input tokens 2026-08-24
gpt-5.6-luna $0.20 $1.20 above 272,000 input tokens 2026-08-24
gemini-3.7-flash $0.75 $3.75 none 2026-08-24
gemini-2.5-flash $0.30 $2.50 none 2026-08-24

None of these is offered as current or authoritative. Each vendor's own pricing page is listed in the sources below, and it is the number to trust over ours.

The same dated preset prices cost per successful task and the token counter.

What this does not count

Five real costs that sit outside the total
Not in the totalWhat to do about it
Batch, off-peak, and cached-input rates The common way a judge actually runs. This models no discount, so halve the rate by hand before you type it.
Retries and failed calls Add a margin of your own. A judge that fails one call in fifty costs two percent more than this.
Drift in judge output length Re-read the token counts against a real judge transcript rather than an estimate.
The cost of building the set That belongs to the set, not to the run. You pay it once and run against it many times.
Pairwise grading This page prices pointwise judging. For pairwise, judge calls per case are k - 1 against an anchor or k(k - 1)/2 for a round robin, per trial. Work it out and enter that number.

Questions

Why is the human column so much larger than the machine column?

Because a person costs dollars a minute and a judge call costs fractions of a cent. In the worked example above, reviewing one item in ten accounts for 93.5 percent of the bill. The two levers that move it are the review share and the minutes per review, and both are worth measuring rather than guessing: time ten real reviews before you type a number here.

What counts as a judge call?

One call to whatever scores an output: an LLM judge, a rubric grader, a classifier. If you score each output on three rubric dimensions with three separate calls, that is three judge calls per graded item. If one call returns all three scores, it is one. Vendors name this differently, grader and evaluator and scorer, and they all bill the same way.

Where do I get a cost per judge call?

Either measure it or build it. Measuring it is better: run fifty judge calls and divide the billed amount by fifty, which captures retries and the prompt you actually send. Building it means switching this page to per-token mode and entering the tokens in one call, at which point the arithmetic is the same one your provider does.

Are these prices current?

They were read from each vendor's own page on the date shown beside every rate, and the table above shows all of them. They are offered as a convenience, not as an authority, and a price can change the day after it is read. Every stamp on this page names the page it came from so you can check it in one click. Your own number overrides ours everywhere.

Does the plan get cheaper per case as it grows?

No. There is no fixed cost in this model, so cost per case is flat: a plan twice the size costs twice as much and the per-case figure does not move. Real evals do have fixed costs, mostly in building the set and in engineering time, and those sit outside this total on purpose so the linearity stays a property you can reason about.

Can I price a pairwise comparison here?

Indirectly. This page prices pointwise judging, one judgment per graded item. For pairwise, work out how many comparisons your design needs per case, anchored against a baseline or round robin across every pair, and enter that as the judge calls per graded item. The arithmetic then holds.

Sources

  1. Anthropic model pricing: the primary record of the Claude rates this page offers as presets, and of the note that Claude Sonnet 5's introductory price is now the standard price with the announced September 2026 increase canceled. Price provenance, not method authorityAnthropic Retrieved
  2. OpenAI model pricing: the primary record of the GPT-5.6 rates this page offers as presets, quoted per 1M tokens with separate short-context and long-context columns. Price provenance, not method authorityOpenAI Retrieved
  3. GPT-5.6 luna model card: the primary record of the long-context rule this page applies, that prompts above 272K input tokens are priced at 2x input and 1.5x output for the full request. Price provenance, not method authorityOpenAI Retrieved
  4. GPT-5.6 terra model card: the same 272K threshold and the same multipliers, which is why both OpenAI presets carry a long-context tier and the Anthropic and Gemini presets do not. Price provenance, not method authorityOpenAI Retrieved
  5. Gemini API pricing: the primary record of the Gemini rates this page offers, including the dated step where Gemini 3.7 Flash moves from $0.75 and $3.75 per 1M tokens to $1.50 and $7.50 on 1 January 2027. Price provenance, not method authorityGoogle Retrieved
  6. Kim, Augmenting Human Evaluation with LLM Judges: How Many Human Reviews Do You Need? A two-stage sampling design where an LLM scores everything and humans score a subsample, sized against a precision target rather than a budget. The principled version of the review-share input this page takes as a numberarXiv Retrieved