Cost per case
not priced
INSTRUMENT | Eval statistics
6 cited sources
Price an eval plan before you run it: cases by variants by trials by judge calls, plus the human review hours it needs, from your own rates.
An eval plan is a multiplication before it is a budget. This prices one from your own rates, on the evidence in our work on how many runs an eval needs and the statistics behind them.
Size it first: cases is the distinct-item count, trials the repeats per case, variants the number of arms. The sample size and power calculator sizes the set. Human review time is usually the largest line.
Showing your last valid result. Update the inputs above to recompute.
Graded items
3,000
3,000 graded items and 6,000 judge calls at these counts. Enter your rates to price this plan.
Cost per case
not priced
Cost per graded item
not priced
These are your own rates. We filled nothing in.
Enter your rates to price this plan.
| Line | Amount |
|---|---|
| Graded items | 3,000 |
| Judge calls | 6,000 |
| Judge spend | not priced |
| Run spend | not priced |
| Machine spend | not priced |
| Human review hours | not priced |
| Human spend | not priced |
| Total spend | not priced |
What the same plan costs at other trial counts
| Trials | Graded items | Machine | Human | Total spend |
|---|---|---|---|---|
| 1 | 600 | not priced | not priced | not priced |
| 2 | 1,200 | not priced | not priced | not priced |
| 3 | 1,800 | not priced | not priced | not priced |
| 5 | 3,000 | not priced | not priced | not priced |
| 10 | 6,000 | not priced | not priced | not priced |
total = judge calls · price + items · run cost + hours · rateHow?
The multiplication. Cases times variants times trials gives the graded items, and each graded item takes however many judge calls you set. Nothing here is a fixed cost, so the total is exactly linear in all three counts: double the trials and you double the bill. That is the property the sensitivity table below the result reads off, and it is why deciding the counts is the budget decision. Sizing them is a question of statistical power, not of taste, and it belongs upstream of this page.
The worked example, small enough to check by hand. 200 cases, 3 variants and 5 trials is 3,000 graded items and 6,000 judge calls. At $0.004 a call plus a $0.0025 platform fee, judging costs $39.00; at $0.02 to produce each output, the run costs $60.00. Reviewing 10 percent of 3,000 items at 3 minutes each is 15.0 hours, which at $95 an hour is $1,425.00. Total $1,524.00, of which the human column is 93.5 percent. That last number is the one most plans get wrong.
Two ways to price the judge. Per call is what you use when you already know the blended cost of one judgment. Per token is what you use when you know the prompt: input tokens times the input price plus output tokens times the output price, divided by a million. Some vendors re-price a long request entirely, so a call above their threshold costs more per token across the whole call and not just above the line. Where a preset carries that rule, the page applies it and the stamp says so.
The rates are yours. Every money field starts empty. The presets are prices a person read from the vendor's own page on the date shown beside each one, offered because typing four of them is tedious, and they can be out of date the day after they were read. Your own number always wins, and editing a filled field turns the stamp back to your number.
Rounding. Everything computes at full precision and rounds only for display. Money shows two decimals, or four when a per-item figure is under a dollar, because a real sub-cent cost printed as $0.00 is worse than useless. Review time is never rounded on its way to hours: 10 percent of 3,005 items is 300.5 items and 15.025 hours. Full precision has one outer limit worth stating: totals are exact to the cent up to about $90 trillion, which is where double-precision arithmetic runs out of cents. The maxima on this form can be combined into a plan past that line, and the cents in such a total are approximate. No budget anyone brings here is within four orders of magnitude of it.
Honest limits. This is arithmetic on numbers you supply. It has no opinion on whether your plan is big enough to answer your question, and it does not know what your judge agrees with. How much human review to buy has a principled answer, sized against a power target rather than a budget, and it is a different calculation from this one. Treat the total as a floor: the lines under "What this does not count" below are all real money and none of them are in it.
Formula: total = judge calls · price + items · run cost + hours · rate
Per 1M tokens in US dollars, by the vendor's own model id. Above the threshold in the long-context column, the whole call bills at 2x input and 1.5x output. Gemini 3.7 Flash is due for a re-read by 2026-12-31
| Model | Input | Output | Long-context rule | We read it |
|---|---|---|---|---|
| claude-sonnet-5 | $2.00 | $10.00 | none | 2026-08-24 |
| claude-haiku-4-5 | $1.00 | $5.00 | none | 2026-08-24 |
| claude-opus-5 | $5.00 | $25.00 | none | 2026-08-24 |
| gpt-5.6-terra | $2.00 | $12.00 | above 272,000 input tokens | 2026-08-24 |
| gpt-5.6-luna | $0.20 | $1.20 | above 272,000 input tokens | 2026-08-24 |
| gemini-3.7-flash | $0.75 | $3.75 | none | 2026-08-24 |
| gemini-2.5-flash | $0.30 | $2.50 | none | 2026-08-24 |
None of these is offered as current or authoritative. Each vendor's own pricing page is listed in the sources below, and it is the number to trust over ours.
The same dated preset prices cost per successful task and the token counter.
| Not in the total | What to do about it |
|---|---|
| Batch, off-peak, and cached-input rates | The common way a judge actually runs. This models no discount, so halve the rate by hand before you type it. |
| Retries and failed calls | Add a margin of your own. A judge that fails one call in fifty costs two percent more than this. |
| Drift in judge output length | Re-read the token counts against a real judge transcript rather than an estimate. |
| The cost of building the set | That belongs to the set, not to the run. You pay it once and run against it many times. |
| Pairwise grading | This page prices pointwise judging. For pairwise, judge calls per case are k - 1 against an anchor or k(k - 1)/2 for a round robin, per trial. Work it out and enter that number. |
Because a person costs dollars a minute and a judge call costs fractions of a cent. In the worked example above, reviewing one item in ten accounts for 93.5 percent of the bill. The two levers that move it are the review share and the minutes per review, and both are worth measuring rather than guessing: time ten real reviews before you type a number here.
One call to whatever scores an output: an LLM judge, a rubric grader, a classifier. If you score each output on three rubric dimensions with three separate calls, that is three judge calls per graded item. If one call returns all three scores, it is one. Vendors name this differently, grader and evaluator and scorer, and they all bill the same way.
Either measure it or build it. Measuring it is better: run fifty judge calls and divide the billed amount by fifty, which captures retries and the prompt you actually send. Building it means switching this page to per-token mode and entering the tokens in one call, at which point the arithmetic is the same one your provider does.
They were read from each vendor's own page on the date shown beside every rate, and the table above shows all of them. They are offered as a convenience, not as an authority, and a price can change the day after it is read. Every stamp on this page names the page it came from so you can check it in one click. Your own number overrides ours everywhere.
No. There is no fixed cost in this model, so cost per case is flat: a plan twice the size costs twice as much and the per-case figure does not move. Real evals do have fixed costs, mostly in building the set and in engineering time, and those sit outside this total on purpose so the linearity stays a property you can reason about.
Indirectly. This page prices pointwise judging, one judgment per graded item. For pairwise, work out how many comparisons your design needs per case, anchored against a baseline or round robin across every pair, and enter that as the judge calls per graded item. The arithmetic then holds.