LatentEval
Directory

Evaluation and reliability calculators

45 calculators | 10 topics

Calculators for AI agent eval statistics: confidence intervals, paired significance, repeated-run reliability, judge calibration, agreement and bias, prompt robustness, and RAG.

How to read these

The instruments

Instrument | Reliability testing

Acceptable Error Rate Calculator: The Reliability Bar Automation Needs

Invert your own manual cost, automation cost, and consequence cost into the minimum reliability that makes automating pay, and see if a measured rate clears it.

Instrument | Reliability testing

Eval Run Register: Preregistration and Run Cards for LLM Evals

Keep one row per eval run in this browser: freeze the plan before results exist, log what happened, and export a run card or a preregistration record for a PR.

Instrument | Propagation & containment

Expected Cost of Failure Calculator for AI Agents

Price a failure rate in money: what gets caught, what escapes, and what correction and rework savings are worth, split by failure class.

Instrument | Evaluation and reliability calculators

Reproducibility Checklist and Manifest Validator

Check whether your write-up records what someone else would need to rerun your eval: a checklist against the published card schema, and a validator for a pasted or uploaded manifest.

Instrument | Propagation & containment

Agent Action Risk Matrix: Which Actions Need an Approval Gate

Place every action your agent can take on two axes: whether anything undoes it, and its blast radius, the field's name for propagation radius. Each row returns a written control level, never a score.

Instrument | Eval statistics

Paired vs Independent: Choosing an Eval Comparison Design

Six questions give you the design, the test that fits it, the rules that fired, and what the branch you did not pick costs in cases and model runs.

Instrument | Eval statistics

pass^k, pass@k and reliability@k Estimator: Chance That All k Runs Pass

Estimate per-task pass^k and suite-level reliability@k from repeated eval runs, with Wilson and Student-t confidence intervals and the unbiased pass@k capability counterpart.

Instrument | Reliability testing

Irreversible Action Inventory: What Your Agent Cannot Take Back

List every action your agent can take, mark what reverses it, and name the point of no return. A written verdict per row, never a score or a budget.

Instrument | Eval statistics

Golden Set Size Planner: How Many Examples an Eval Set Needs

Size a golden set from a baseline pass rate, a margin, and the drop worth catching: items per slice, a total, and a verdict on the set you already have.

Instrument | Eval statistics

Wilson and Clopper-Pearson Pass-Rate Confidence Interval Calculator

Turn an eval pass rate (k of n) into a defensible confidence interval: Wilson score, Clopper-Pearson exact, and Wald normal bounds side by side.

Instrument | Eval statistics

Repeated-Run Variance Planner: How Many Trials per Eval Case

Split your eval's spread into run-to-run and case-to-case variance, then read the trials per case, the case count that implies, and the protocol to freeze.

Instrument | Eval statistics

Eval Sample Size and Power Calculator

Find how many eval runs you need to detect a pass-rate drop at a target power, across two-arm, fixed-baseline, and paired McNemar designs.

Instrument | Evaluation and reliability calculators

Judge Validation Report Builder: Report, Version Card, Drift Plan

Fill one form and keep three documents: a validation report showing your LLM judge was checked against people, a version card pinning the prompt and settings, and a drift plan with a date on it.

Instrument | Eval statistics

Eval Budget Calculator: What an LLM Evaluation Run Will Cost

Price an eval plan before you run it: cases by variants by trials by judge calls, plus the human review hours it needs, from your own rates.

Instrument | Eval statistics

McNemar Test Calculator for Paired Eval Runs

Run McNemar's paired significance test on two eval runs scored on the same items, with exact and chi-square p-values and a CI on the pass-rate difference.

Instrument | Propagation & containment

Multi-Step Agent Reliability Calculator (Independent Steps)

Compute the end-to-end success rate of a multi-step agent chain from its per-step reliability, and the per-step budget your target reliability demands.

Instrument | Judge reliability

Cohen's Kappa Calculator (Two Raters): Inter-Rater Reliability

Cohen's kappa from the four counts of a 2x2 table: chance-corrected agreement between two raters, with a confidence interval and an interpretation band. Built for an LLM judge against human labels.

Instrument | Judge reliability

LLM Judge Bias Correction Calculator

Correct an LLM judge's raw pass rate for its measured sensitivity and specificity (the Rogan-Gladen correction), then report a calibration-aware confidence interval on the true rate.

Instrument | Judge reliability

LLM Judge Position Bias Calculator (Swap-Consistency Test)

Score a pairwise LLM judge on order-swapped pairs: swap-consistency and position-preference rates with Wilson intervals and an exact test for position bias.

Instrument | RAG & retrieval

RAG Chunk Size Calculator: Index Size and Context Budget

Size a RAG index from its chunking plan: chunk size, overlap, and top-k give the vector count, storage inflation, and context-budget share. Which size retrieves best stays a measured question.

Instrument | Reliability testing

LLM Token Counter: Context Budget and Cost per Attempt

Count tokens for a named model, see what the request reserves in the context window once output is set aside, and price one attempt at your own rates. Exact where we bundle the encoding.

Instrument | Judge reliability

LLM Judge Calibration Calculator: Brier Score and ECE

Score an LLM judge's stated confidence against observed correctness: Brier score, expected and maximum calibration error, a reliability diagram, and the error rate behind its confident verdicts.

Instrument | Eval statistics

LLM Eval A/B Comparator: Paired Difference with Bootstrap CIs

Run an LLM eval A/B test on paired graded scores: mean score difference with a bootstrap confidence interval, paired effect size, and win/tie/loss evidence.

Instrument | Judge reliability

Pairwise Win Rate Calculator with Elo and Bradley-Terry Ratings

Turn pairwise LLM judge comparisons between models or prompts into win rates with Wilson intervals, Bradley-Terry strengths on the Elo scale, and a plain read on whether the advantage is real.

Instrument | Reliability testing

Human Review Threshold Optimizer for Confidence-Gated Automation

Derive the human review threshold from your own records: paste stated confidence and observed correctness, price one review against one missed failure, and read the cut that costs least.

Instrument | Judge reliability

Fleiss' Kappa, Krippendorff's Alpha and ICC Calculator

Agreement among three or more raters or LLM judges: Fleiss' kappa, Krippendorff's alpha, Gwet's AC1 and intraclass correlation, with percent agreement, bootstrap intervals, and per-rater diagnostics.

Instrument | Evaluation and reliability calculators

Judge Agreement Tracker: Log Kappa Over Time

Keep a dated log of one judge's agreement with human labels: the value, the prompt version and set behind it, the published band it falls in, the change since the last run, and the next check date.

Instrument | Propagation & containment

Retry, Verification, and Fallback Reliability Optimizer for LLM Agents

How many times should you retry a failed agent step? Compare retry, majority vote, verifier gate, fallback, and escalation on delivered correctness and expected cost per task, from your own rates.

Instrument | Reliability testing

AI SLO Error Budget Calculator: Split a Failure Budget Across Classes

Turn a completion-rate target and a task volume into a failure budget split by class, with a Wilson interval on each class observed rate against its allowance.

Instrument | Eval statistics

Benchmark Rank Uncertainty Calculator: Score Intervals and Rank Ranges

Put error bars on an LLM benchmark leaderboard. Paste scores with item counts, correct-of-total, or standard errors and see which adjacent ranks are a statistical tie.

Instrument | Reliability testing

Prompt Robustness Analyzer: Perturbation Stability for LLM Evals

Is a prompt's pass-rate drop real or noise? Paste pass/fail results for a baseline prompt and its semantically equivalent variants and read each change with a paired interval and an exact McNemar p.

Instrument | RAG & retrieval

RAG Retrieval Quality Calculator: Precision@K, Recall@K, nDCG, MRR, MAP

Score a retrieval run per query and across the suite: Precision@K, Recall@K, Hit@K, nDCG, mean reciprocal rank (MRR), and mean average precision (MAP), with a K sweep and intervals.

Instrument | Propagation & containment

Multi-Step Agent Reliability with Shared Failures (Cascade Simulator)

Price a multi-step agent chain when its steps share a cause: end-to-end reliability under common-cause failure beside the independent figure, plus the retry ceiling that shared failures impose.

Instrument | Eval statistics

Eval Dataset Schema Validator: Golden-Set Checks and Export Profiles

Check an eval golden set against a neutral schema in your browser: findings by row and line, duplicate ids, field completeness, plus a template and exports for four eval tools.

Instrument | Eval statistics

Exact Duplicate Row Checker for Eval and Golden Datasets

Find byte-identical rows in an eval or golden dataset without uploading it. Duplicate groups with row and line numbers, an identity key you choose, and the sample size left after a dedupe.

Instrument | Eval statistics

Near-Duplicate Row Detector: Find Eval Rows That Differ Only in Formatting

Find rows in an eval set that are the same test written twice. Normalize the columns you pick, group rows whose normalized text matches, and read the exact transform, the raw variants, and a CSV.

Instrument | Eval statistics

Cross-Split Overlap Checker: Find Eval Items in Both Train and Test

Compare two eval splits in your browser and see which items sit on both sides, with a separate overlap rate against the row count of each split, exact and normalized matching, and the lines to open.

Instrument | Eval statistics

N-Gram Overlap Checker: Find Contamination Between Two Eval Files

Check a test set against a training file for shared word sequences. Pick n, read the per-row dirty-token share, see the matching text in both files, and export a CSV. Both files stay in your browser.

Instrument | Eval statistics

Slice and Class Balance Analyzer: See What an Eval Set Contains

Read what an eval set holds before you quote a pass rate. Count rows per slice, per class, and per control flag, put an interval on every share, and see the smallest slice named.

Instrument | Propagation & containment

Cost per Successful Task Calculator: What One Acceptable Result Costs

Turn a per-attempt price and a pass rate into the cost of one acceptable result, with retries, failed attempts, review and the fallback path all priced into the total.

Instrument | Evaluation and reliability calculators

Evaluation Card Generator: Markdown and JSON on a Published Schema

Write one evaluation down so a second person can read the number correctly, or run it again. Fill the form, export a Markdown card and a JSON card that validates against a schema published here.

Instrument | Reliability testing

Golden Set Version Tracker: Dataset Diff and Changelog

Diff two versions of your golden eval set in the browser, keep a dated changelog of what changed and why, and freeze the baseline metric values the next run is judged against.

Instrument | Reliability testing

Release Gate Designer: Go/No-Go Rules and Decision Log

Write the rule that turns your eval figures into PROMOTE, HOLD or ROLLBACK, get the verdict plus the rule that produced it, and keep a dated decision log and memo in your browser.

Instrument | Reliability testing

Change Impact Worksheet: What to Rerun Before You Ship

Describe one change to a shipped AI system and get the rerun list it invalidates: which suites, slices and graders to run again, why each is on the list, and a change record for the ticket.

Instrument | Eval statistics

Independent Two-Proportion Test Calculator for Eval Arms

Compare two pass rates measured on different case sets: a pooled two-proportion z test, Fisher exact when an arm is small, and an interval on the difference.