Evaluator
LLM judge
INSTRUMENT | Eval statistics
Answer four questions about oracle type, grader, system count, and failure mode. Get the admissible evaluator, inadmissible ones and why, the metric family, and a route to the tool that runs it.
Answer four questions and get the evaluator type that is admissible for your oracle, the ones that are not and why, the metric family, and which tool runs the number. For the comparison-design question specifically, use the paired vs. independent design chooser. The oracle-evaluator admissibility rules are documented in the evaluation case schema. Built on the significance walkthrough.
A set built from any pack carries several oracle types across its families, so the navigator runs once per family and acceptance criteria are written per family.
The constraint matrix is our own synthesis of the admissibility rules implicit in the evaluation literature. It is ours, not a standard.
| Oracle type | Exact or programmatic check | Rubric-scored human | LLM judge | Pairwise preference | Trajectory or state check | Domain expert |
|---|---|---|---|---|---|---|
| Reference answer | Admissible | Admissible | With conditions | Inadmissible | Inadmissible | Admissible |
| Acceptable set | Admissible | Admissible | With conditions | Inadmissible | Inadmissible | Admissible |
| Rubric | Inadmissible | Admissible | With conditions | Inadmissible | Inadmissible | Admissible |
| Expected tool calls | Admissible | Inadmissible | Inadmissible | Inadmissible | Admissible | Admissible |
| Expected final state | Admissible | Inadmissible | Inadmissible | Inadmissible | Admissible | Admissible |
| Expected behavior only | Inadmissible | Inadmissible | Inadmissible | Inadmissible | Inadmissible | Admissible |
Comparison design x metric family: routing
| Rule | When it fires | Metric | Route |
|---|---|---|---|
| M1 | two systems scored the same cases | Paired difference | paired-vs-independent-design-chooser |
| M2 | two systems scored different cases | Pass rate with an interval | independent-two-proportion-calculator |
| M3 | one system measured alone, and every attempt must succeed | pass^k | reliability-at-k-estimator |
| M4 | one system measured alone, and at least one attempt needs to succeed | pass@k | reliability-at-k-estimator |
| M5 | one system measured alone | Pass rate with an interval | pass-rate-ci-calculator |
Showing your last valid result. Update the inputs above to recompute.
Evaluator
LLM judge
Metric
Pass rate with an interval
Quality is multi-dimensional and written as criteria. With no human graders, a model judge is the admissible automated alternative, with its own agreement evidence.
The share of cases that passed, reported with a confidence interval on the share.
Inadmissible evaluators
| Evaluator | What it is |
|---|---|
| Exact or programmatic check | Code decides. String match after a stated normalization, a parser, a schema check, a set membership test. |
| Pairwise preference | A reader or a judge picks between two outputs. Inadmissible as an oracle-correctness evaluator because it answers which of two outputs is preferred, not which is correct. Where the claim itself is a preference or quality ranking, pairwise is the right instrument. |
| Trajectory or state check | The path or the world after: which tools were called, what state changed, which constraint was violated. |
Quality here is multi-dimensional and written as criteria. A string comparison cannot read a criterion. A trajectory check inspects the path, not the quality of the output.
Conditionally admissible evaluators
| Evaluator | Conditions |
|---|---|
| LLM judge | Admissible only with agreement evidence against human labels on your own items. Route to judge-validation-report-builder. |
The linked tools on this site can run these checks.
| Stage | Rule | What happened | Stopped at |
|---|---|---|---|
| 1 | E1 | Stopped | the oracle is a reference answer |
| 1 | E2 | Stopped | the oracle is an acceptable set |
| 1 | E3 | Fired | |
| 1 | E4 | Stopped | human graders are available on a sample |
| 1 | E5 | Stopped | human graders are available for every case |
| 1 | E6 | Stopped | a domain expert is available |
| 1 | E7 | Stopped | the oracle is expected tool calls |
| 1 | E8 | Stopped | the oracle is the expected final state |
| 1 | E9 | Stopped | the oracle is expected behavior only |
| 2 | M1 | Stopped | two systems scored the same cases |
| 2 | M2 | Stopped | two systems scored different cases |
| 2 | M3 | Stopped | every attempt must succeed |
| 2 | M4 | Stopped | at least one attempt needs to succeed |
| 2 | M5 | Fired |
Two passes, not one. The first pass reads the oracle type and who is available to grade, and returns the evaluator. The second reads how many systems and what happens on failure, and returns the metric family plus a route to the live tool that computes it. Both are first-match rule tables: every rule is checked, the first that matches fires, and the rest are recorded as beaten or stopped.
Pass one: the evaluator
| Rule | When it fires | Evaluator |
|---|---|---|
| E1 | the oracle is a reference answer | exact-or-programmatic |
| E2 | the oracle is an acceptable set | exact-or-programmatic |
| E3 | the oracle is a rubric, and no human graders are available | llm-judge |
| E4 | the oracle is a rubric, and human graders are available on a sample | llm-judge |
| E5 | the oracle is a rubric, and human graders are available for every case | rubric-human |
| E6 | the oracle is a rubric, and a domain expert is available | human-expert |
| E7 | the oracle is expected tool calls | trajectory-or-state-check |
| E8 | the oracle is the expected final state | trajectory-or-state-check |
| E9 | the oracle is expected behavior only | human-expert |
Pass two: the metric
| Rule | When it fires | Metric | Route |
|---|---|---|---|
| M1 | two systems scored the same cases | Paired difference | paired-vs-independent-design-chooser |
| M2 | two systems scored different cases | Pass rate with an interval | independent-two-proportion-calculator |
| M3 | one system measured alone, and every attempt must succeed | pass^k | reliability-at-k-estimator |
| M4 | one system measured alone, and at least one attempt needs to succeed | pass@k | reliability-at-k-estimator |
| M5 | one system measured alone | Pass rate with an interval | pass-rate-ci-calculator |
Admissibility. The constraint matrix maps each oracle type to the evaluators it admits and the ones it rules out. Pairwise preference is inadmissible as an oracle-correctness evaluator for every oracle type here, because it measures which of two outputs a reader prefers, not whether either is correct against the oracle. Where selecting between candidates is the actual question, the pairwise win rate calculator's protocol chooser settles pointwise versus pairwise.
Validity evidence. Some evaluators need their own proof that the grading is sound. A model judge needs agreement against human labels on your own items, position-sensitivity, length-bias, and self-preference checks. A rubric-scored human evaluator needs inter-rater agreement corrected for chance. A domain expert needs at least two experts agreeing on a sample. An exact-or-programmatic evaluator needs none: the code is the definition.
Pairwise preference measures which of two outputs a reader prefers, not whether either is correct against the oracle. It answers a different question. Where selecting between candidates is the actual question, the pairwise win rate calculator's protocol chooser is the right tool.
This happens when the oracle is expected behavior only and no human grader is available. The path forward is to turn the behavior description into a rubric, an acceptable set, or a set of expected calls, which opens the door to other evaluators.
pass^k, pass@k, and plain pass rate answer different reliability questions. pass^k (every attempt must succeed) falls with k; pass@k (at least one must succeed) rises with k. The retry model picks the one that matches what the product actually experiences.
The trace shows every rule the engine checked, which one fired, and where each non-firing rule stopped. It is the audit trail: if the recommendation looks wrong, the trace tells you which clause made it so.
No. It names the evaluator type and the metric family and links to the tool that runs the number. Each route in the recommendation block is a link to the next step.