LatentEval
Topic

Evaluation design calculators

11 calculators | 5 analyses

Everything that happens before a run: what the system is for, which failures matter, what a case looks like, who grades it, and what the result will and will not prove.

The method

What these numbers mean

Most of what makes an eval wrong is decided before anything runs. Which population the cases come from, whether the expected answer is a reference or a judgment, which failures get their own count, and which question the number will be asked to settle are all design choices. A confidence interval computed afterward cannot recover any of them. The instruments below are for that stage. None of them runs a suite and none of them calls a model.

What you are trying to work outThe instrument
What this system is actually for, and what that changes about testing itprofile the context of use
What the whole evaluation will test, written down before results existwrite the evaluation blueprint
Which test cases a named failure concern actually calls formap concerns to test cases
Which parts of the testing surface have no cases at allbuild the coverage matrix
Which slices to plan for, and how many cases each one needsplan the slices before the set exists
Whether to score both systems on the same cases or different oneschoose paired or independent
How many times to run each case before the spread is readableplan the repeated runs
What the design will cost before you commit to itestimate the eval budget
Whether to score outputs one at a time or against each otherchoose the scoring protocol
Who or what should grade this, and which metric answers the questionchoose the evaluator and the metric
What the evidence you collected actually licenses you to claimcheck what the evaluation proves
What a case looks like for the task family you are testingstart from a task-family blueprint
How many distinct items the set needs, per slice and in totalsize the golden set
What "good enough to ship" means, in writing, before the runset the acceptance criteria

This family rests on our own work on the four ways an eval number lies, which eval number needs which test, and what a single pass rate hides about run-to-run consistency.

The fork people get stuck on is whether to write more cases or to write different ones. Coverage of the failures you already named is a different property from coverage of the population you actually serve, and only the first is visible in a grid. The coverage matrix shows that. The context profile tells you whether the second one is the property you should be worried about.

The second fork is the expected-answer column. A set that mixes a reference answer, a rubric and an expected end state in one column and then applies one grader to all three produces a number whose meaning changes row by row. Declaring the oracle type per case is what the case schema is for, and the evaluator navigator lists which graders the declared oracle admits, and the schema validator checks each row against that list.

In this cluster

Tools in this topic

Instrument | Evaluation design

Context-of-Use Profiler: Turn a System Description into Evaluation Consequences

Describe what your AI system does, who it affects and how it fails, and get the failure families, case classes, slices and evidence expectations those answers demand.

Instrument | Evaluation design

Evaluation Blueprint Builder: Write the Argument Before You Run the Eval

State each claim your evaluation will make, name the evidence that would support or refute it, and write the inference step connecting the two, all before any results exist.

Instrument | Evaluation design

Risk-to-Test Mapper: Which Failures Need Which Cases

Name the failures you care about, then see the test-case shapes, slices, evaluators and metrics each one needs. Every mapping says what it cannot establish. The mapping is a LatentEval synthesis.

Instrument | Evaluation design

Evaluator and Metric Navigator: Which Grader, Which Number

Answer four questions about oracle type, grader, system count, and failure mode. Get the admissible evaluator, inadmissible ones and why, the metric family, and a route to the tool that runs it.

Instrument | Evaluation design

What Does This Evaluation Prove?

Select the evidence your evaluation produced, see which claims it supports, which it cannot, and the inference error behind each gap. Each unsupported claim names the instrument that would close it.

Instrument | Evaluation design

Judge Validation Report Builder: Report, Version Card, Drift Plan

Fill one form and keep three documents: a validation report showing your LLM judge was checked against people, a version card pinning the prompt and settings, and a drift plan with a date on it.

Instrument | Evaluation design

Release Gate Designer: Go/No-Go Rules and Decision Log

Write the rule that turns your eval figures into PROMOTE, HOLD or ROLLBACK, get the verdict plus the rule that produced it, and keep a dated decision log and memo in your browser.

Instrument | Evaluation design

Evaluation Coverage Matrix: Which Cells Have No Cases

Put what you test on one axis and what you test it across on the other. The grid shows which cells have cases, which are empty, which are too small to carry an interval, and what each gap costs you.

Instrument | Evaluation design

Acceptance Criteria Builder: Set the Bar Before the Run

State each pass/fail criterion, the risk axes it covers, the reasoning behind the threshold and who set it. Get the sample size each one needs and export the rules for the release gate designer.

Instrument | Evaluation design

Golden Set Blueprint Builder: Seed a Test Set From a Task Family

Choose a task family, load its case taxonomy and starter rows, then edit, add and export. Every row carries an oracle type and a provenance tag, and the export validates against the eval-set schema.

Instrument | Evaluation design

Claim-Evidence Matrix: What Backs Each Deployment Claim

One row per deployment claim, eleven evidence dimensions from relevance and representativeness to independence and applicability, with named rate-down reasons and residual deficits.