LatentEval

INSTRUMENT | Eval statistics

Evaluator and Metric Navigator: Which Grader, Which Number

Answer four questions about oracle type, grader, system count, and failure mode. Get the admissible evaluator, inadmissible ones and why, the metric family, and a route to the tool that runs it.

Answer four questions and get the evaluator type that is admissible for your oracle, the ones that are not and why, the metric family, and which tool runs the number. For the comparison-design question specifically, use the paired vs. independent design chooser. The oracle-evaluator admissibility rules are documented in the evaluation case schema. Built on the significance walkthrough.

A set built from any pack carries several oracle types across its families, so the navigator runs once per family and acceptance criteria are written per family.

The constraint matrix is our own synthesis of the admissibility rules implicit in the evaluation literature. It is ours, not a standard.

What form of correct answer do you have?

Check this value.

Who is available to grade?

Check this value.

How many systems are you measuring?

Check this value.

Oracle type x evaluator type: admissibility
Oracle type Exact or programmatic checkRubric-scored humanLLM judgePairwise preferenceTrajectory or state checkDomain expert
Reference answer Admissible Admissible With conditions Inadmissible Inadmissible Admissible
Acceptable set Admissible Admissible With conditions Inadmissible Inadmissible Admissible
Rubric Inadmissible Admissible With conditions Inadmissible Inadmissible Admissible
Expected tool calls Admissible Inadmissible Inadmissible Inadmissible Admissible Admissible
Expected final state Admissible Inadmissible Inadmissible Inadmissible Admissible Admissible
Expected behavior only Inadmissible Inadmissible Inadmissible Inadmissible Inadmissible Admissible

Comparison design x metric family: routing

RuleWhen it firesMetricRoute
M1 two systems scored the same cases Paired difference paired-vs-independent-design-chooser
M2 two systems scored different cases Pass rate with an interval independent-two-proportion-calculator
M3 one system measured alone, and every attempt must succeed pass^k reliability-at-k-estimator
M4 one system measured alone, and at least one attempt needs to succeed pass@k reliability-at-k-estimator
M5 one system measured alone Pass rate with an interval pass-rate-ci-calculator

Evaluator

LLM judge

Metric

Pass rate with an interval

Quality is multi-dimensional and written as criteria. With no human graders, a model judge is the admissible automated alternative, with its own agreement evidence.

The share of cases that passed, reported with a confidence interval on the share.

Evaluators that are not admissible for this oracle type

Inadmissible evaluators

EvaluatorWhat it is
Exact or programmatic check Code decides. String match after a stated normalization, a parser, a schema check, a set membership test.
Pairwise preference A reader or a judge picks between two outputs. Inadmissible as an oracle-correctness evaluator because it answers which of two outputs is preferred, not which is correct. Where the claim itself is a preference or quality ranking, pairwise is the right instrument.
Trajectory or state check The path or the world after: which tools were called, what state changed, which constraint was violated.

Quality here is multi-dimensional and written as criteria. A string comparison cannot read a criterion. A trajectory check inspects the path, not the quality of the output.

Evaluators admissible with conditions

Conditionally admissible evaluators

EvaluatorConditions
LLM judge Admissible only with agreement evidence against human labels on your own items. Route to judge-validation-report-builder.
Validity evidence the chosen evaluator needs
  • Agreement against human labels: Measure the judge against human labels on your own items, not a published benchmark.
  • Position sensitivity: Swap the order of candidates and check whether the verdict changes.
  • Length bias: Check whether the judge favors longer outputs independently of quality.
  • Self-preference: Check whether the model prefers its own outputs over equally good alternatives.

The linked tools on this site can run these checks.

Every rule we checked, and where each one stopped
StageRuleWhat happenedStopped at
1 E1 Stopped the oracle is a reference answer
1 E2 Stopped the oracle is an acceptable set
1 E3 Fired
1 E4 Stopped human graders are available on a sample
1 E5 Stopped human graders are available for every case
1 E6 Stopped a domain expert is available
1 E7 Stopped the oracle is expected tool calls
1 E8 Stopped the oracle is the expected final state
1 E9 Stopped the oracle is expected behavior only
2 M1 Stopped two systems scored the same cases
2 M2 Stopped two systems scored different cases
2 M3 Stopped every attempt must succeed
2 M4 Stopped at least one attempt needs to succeed
2 M5 Fired
How this is calculated

Two passes, not one. The first pass reads the oracle type and who is available to grade, and returns the evaluator. The second reads how many systems and what happens on failure, and returns the metric family plus a route to the live tool that computes it. Both are first-match rule tables: every rule is checked, the first that matches fires, and the rest are recorded as beaten or stopped.

Pass one: the evaluator

RuleWhen it firesEvaluator
E1 the oracle is a reference answer exact-or-programmatic
E2 the oracle is an acceptable set exact-or-programmatic
E3 the oracle is a rubric, and no human graders are available llm-judge
E4 the oracle is a rubric, and human graders are available on a sample llm-judge
E5 the oracle is a rubric, and human graders are available for every case rubric-human
E6 the oracle is a rubric, and a domain expert is available human-expert
E7 the oracle is expected tool calls trajectory-or-state-check
E8 the oracle is the expected final state trajectory-or-state-check
E9 the oracle is expected behavior only human-expert

Pass two: the metric

RuleWhen it firesMetricRoute
M1 two systems scored the same cases Paired difference paired-vs-independent-design-chooser
M2 two systems scored different cases Pass rate with an interval independent-two-proportion-calculator
M3 one system measured alone, and every attempt must succeed pass^k reliability-at-k-estimator
M4 one system measured alone, and at least one attempt needs to succeed pass@k reliability-at-k-estimator
M5 one system measured alone Pass rate with an interval pass-rate-ci-calculator

Admissibility. The constraint matrix maps each oracle type to the evaluators it admits and the ones it rules out. Pairwise preference is inadmissible as an oracle-correctness evaluator for every oracle type here, because it measures which of two outputs a reader prefers, not whether either is correct against the oracle. Where selecting between candidates is the actual question, the pairwise win rate calculator's protocol chooser settles pointwise versus pairwise.

Validity evidence. Some evaluators need their own proof that the grading is sound. A model judge needs agreement against human labels on your own items, position-sensitivity, length-bias, and self-preference checks. A rubric-scored human evaluator needs inter-rater agreement corrected for chance. A domain expert needs at least two experts agreeing on a sample. An exact-or-programmatic evaluator needs none: the code is the definition.

Questions

Questions

Why is pairwise preference inadmissible?

Pairwise preference measures which of two outputs a reader prefers, not whether either is correct against the oracle. It answers a different question. Where selecting between candidates is the actual question, the pairwise win rate calculator's protocol chooser is the right tool.

What if no evaluator is admissible?

This happens when the oracle is expected behavior only and no human grader is available. The path forward is to turn the behavior description into a rubric, an acceptable set, or a set of expected calls, which opens the door to other evaluators.

Why does the metric depend on the retry model?

pass^k, pass@k, and plain pass rate answer different reliability questions. pass^k (every attempt must succeed) falls with k; pass@k (at least one must succeed) rises with k. The retry model picks the one that matches what the product actually experiences.

What is the trace table for?

The trace shows every rule the engine checked, which one fired, and where each non-firing rule stopped. It is the audit trail: if the recommendation looks wrong, the trace tells you which clause made it so.

Does this tool run the metric?

No. It names the evaluator type and the metric family and links to the tool that runs the number. Each route in the recommendation block is a link to the next step.