Evaluation and reliability calculators
45 calculators | 10 topics
Calculators for AI agent eval statistics: confidence intervals, paired significance, repeated-run reliability, judge calibration, agreement and bias, prompt robustness, and RAG.
How to read these
The instruments
Acceptable Error Rate Calculator: The Reliability Bar Automation Needs
Invert your own manual cost, automation cost, and consequence cost into the minimum reliability that makes automating pay, and see if a measured rate clears it.
Instrument | Reliability testingEval Run Register: Preregistration and Run Cards for LLM Evals
Keep one row per eval run in this browser: freeze the plan before results exist, log what happened, and export a run card or a preregistration record for a PR.
Instrument | Propagation & containmentExpected Cost of Failure Calculator for AI Agents
Price a failure rate in money: what gets caught, what escapes, and what correction and rework savings are worth, split by failure class.
Instrument | Evaluation and reliability calculatorsReproducibility Checklist and Manifest Validator
Check whether your write-up records what someone else would need to rerun your eval: a checklist against the published card schema, and a validator for a pasted or uploaded manifest.
Instrument | Propagation & containmentAgent Action Risk Matrix: Which Actions Need an Approval Gate
Place every action your agent can take on two axes: whether anything undoes it, and its blast radius, the field's name for propagation radius. Each row returns a written control level, never a score.
Instrument | Eval statisticsPaired vs Independent: Choosing an Eval Comparison Design
Six questions give you the design, the test that fits it, the rules that fired, and what the branch you did not pick costs in cases and model runs.
Instrument | Eval statisticspass^k, pass@k and reliability@k Estimator: Chance That All k Runs Pass
Estimate per-task pass^k and suite-level reliability@k from repeated eval runs, with Wilson and Student-t confidence intervals and the unbiased pass@k capability counterpart.
Instrument | Reliability testingIrreversible Action Inventory: What Your Agent Cannot Take Back
List every action your agent can take, mark what reverses it, and name the point of no return. A written verdict per row, never a score or a budget.
Instrument | Eval statisticsGolden Set Size Planner: How Many Examples an Eval Set Needs
Size a golden set from a baseline pass rate, a margin, and the drop worth catching: items per slice, a total, and a verdict on the set you already have.
Instrument | Eval statisticsWilson and Clopper-Pearson Pass-Rate Confidence Interval Calculator
Turn an eval pass rate (k of n) into a defensible confidence interval: Wilson score, Clopper-Pearson exact, and Wald normal bounds side by side.
Instrument | Eval statisticsRepeated-Run Variance Planner: How Many Trials per Eval Case
Split your eval's spread into run-to-run and case-to-case variance, then read the trials per case, the case count that implies, and the protocol to freeze.
Instrument | Eval statisticsEval Sample Size and Power Calculator
Find how many eval runs you need to detect a pass-rate drop at a target power, across two-arm, fixed-baseline, and paired McNemar designs.
Instrument | Evaluation and reliability calculatorsJudge Validation Report Builder: Report, Version Card, Drift Plan
Fill one form and keep three documents: a validation report showing your LLM judge was checked against people, a version card pinning the prompt and settings, and a drift plan with a date on it.
Instrument | Eval statisticsEval Budget Calculator: What an LLM Evaluation Run Will Cost
Price an eval plan before you run it: cases by variants by trials by judge calls, plus the human review hours it needs, from your own rates.
Instrument | Eval statisticsMcNemar Test Calculator for Paired Eval Runs
Run McNemar's paired significance test on two eval runs scored on the same items, with exact and chi-square p-values and a CI on the pass-rate difference.
Instrument | Propagation & containmentMulti-Step Agent Reliability Calculator (Independent Steps)
Compute the end-to-end success rate of a multi-step agent chain from its per-step reliability, and the per-step budget your target reliability demands.
Instrument | Judge reliabilityCohen's Kappa Calculator (Two Raters): Inter-Rater Reliability
Cohen's kappa from the four counts of a 2x2 table: chance-corrected agreement between two raters, with a confidence interval and an interpretation band. Built for an LLM judge against human labels.
Instrument | Judge reliabilityLLM Judge Bias Correction Calculator
Correct an LLM judge's raw pass rate for its measured sensitivity and specificity (the Rogan-Gladen correction), then report a calibration-aware confidence interval on the true rate.
Instrument | Judge reliabilityLLM Judge Position Bias Calculator (Swap-Consistency Test)
Score a pairwise LLM judge on order-swapped pairs: swap-consistency and position-preference rates with Wilson intervals and an exact test for position bias.
Instrument | RAG & retrievalRAG Chunk Size Calculator: Index Size and Context Budget
Size a RAG index from its chunking plan: chunk size, overlap, and top-k give the vector count, storage inflation, and context-budget share. Which size retrieves best stays a measured question.
Instrument | Reliability testingLLM Token Counter: Context Budget and Cost per Attempt
Count tokens for a named model, see what the request reserves in the context window once output is set aside, and price one attempt at your own rates. Exact where we bundle the encoding.
Instrument | Judge reliabilityLLM Judge Calibration Calculator: Brier Score and ECE
Score an LLM judge's stated confidence against observed correctness: Brier score, expected and maximum calibration error, a reliability diagram, and the error rate behind its confident verdicts.
Instrument | Eval statisticsLLM Eval A/B Comparator: Paired Difference with Bootstrap CIs
Run an LLM eval A/B test on paired graded scores: mean score difference with a bootstrap confidence interval, paired effect size, and win/tie/loss evidence.
Instrument | Judge reliabilityPairwise Win Rate Calculator with Elo and Bradley-Terry Ratings
Turn pairwise LLM judge comparisons between models or prompts into win rates with Wilson intervals, Bradley-Terry strengths on the Elo scale, and a plain read on whether the advantage is real.
Instrument | Reliability testingHuman Review Threshold Optimizer for Confidence-Gated Automation
Derive the human review threshold from your own records: paste stated confidence and observed correctness, price one review against one missed failure, and read the cut that costs least.
Instrument | Judge reliabilityFleiss' Kappa, Krippendorff's Alpha and ICC Calculator
Agreement among three or more raters or LLM judges: Fleiss' kappa, Krippendorff's alpha, Gwet's AC1 and intraclass correlation, with percent agreement, bootstrap intervals, and per-rater diagnostics.
Instrument | Evaluation and reliability calculatorsJudge Agreement Tracker: Log Kappa Over Time
Keep a dated log of one judge's agreement with human labels: the value, the prompt version and set behind it, the published band it falls in, the change since the last run, and the next check date.
Instrument | Propagation & containmentRetry, Verification, and Fallback Reliability Optimizer for LLM Agents
How many times should you retry a failed agent step? Compare retry, majority vote, verifier gate, fallback, and escalation on delivered correctness and expected cost per task, from your own rates.
Instrument | Reliability testingAI SLO Error Budget Calculator: Split a Failure Budget Across Classes
Turn a completion-rate target and a task volume into a failure budget split by class, with a Wilson interval on each class observed rate against its allowance.
Instrument | Eval statisticsBenchmark Rank Uncertainty Calculator: Score Intervals and Rank Ranges
Put error bars on an LLM benchmark leaderboard. Paste scores with item counts, correct-of-total, or standard errors and see which adjacent ranks are a statistical tie.
Instrument | Reliability testingPrompt Robustness Analyzer: Perturbation Stability for LLM Evals
Is a prompt's pass-rate drop real or noise? Paste pass/fail results for a baseline prompt and its semantically equivalent variants and read each change with a paired interval and an exact McNemar p.
Instrument | RAG & retrievalRAG Retrieval Quality Calculator: Precision@K, Recall@K, nDCG, MRR, MAP
Score a retrieval run per query and across the suite: Precision@K, Recall@K, Hit@K, nDCG, mean reciprocal rank (MRR), and mean average precision (MAP), with a K sweep and intervals.
Instrument | Propagation & containmentMulti-Step Agent Reliability with Shared Failures (Cascade Simulator)
Price a multi-step agent chain when its steps share a cause: end-to-end reliability under common-cause failure beside the independent figure, plus the retry ceiling that shared failures impose.
Instrument | Eval statisticsEval Dataset Schema Validator: Golden-Set Checks and Export Profiles
Check an eval golden set against a neutral schema in your browser: findings by row and line, duplicate ids, field completeness, plus a template and exports for four eval tools.
Instrument | Eval statisticsExact Duplicate Row Checker for Eval and Golden Datasets
Find byte-identical rows in an eval or golden dataset without uploading it. Duplicate groups with row and line numbers, an identity key you choose, and the sample size left after a dedupe.
Instrument | Eval statisticsNear-Duplicate Row Detector: Find Eval Rows That Differ Only in Formatting
Find rows in an eval set that are the same test written twice. Normalize the columns you pick, group rows whose normalized text matches, and read the exact transform, the raw variants, and a CSV.
Instrument | Eval statisticsCross-Split Overlap Checker: Find Eval Items in Both Train and Test
Compare two eval splits in your browser and see which items sit on both sides, with a separate overlap rate against the row count of each split, exact and normalized matching, and the lines to open.
Instrument | Eval statisticsN-Gram Overlap Checker: Find Contamination Between Two Eval Files
Check a test set against a training file for shared word sequences. Pick n, read the per-row dirty-token share, see the matching text in both files, and export a CSV. Both files stay in your browser.
Instrument | Eval statisticsSlice and Class Balance Analyzer: See What an Eval Set Contains
Read what an eval set holds before you quote a pass rate. Count rows per slice, per class, and per control flag, put an interval on every share, and see the smallest slice named.
Instrument | Propagation & containmentCost per Successful Task Calculator: What One Acceptable Result Costs
Turn a per-attempt price and a pass rate into the cost of one acceptable result, with retries, failed attempts, review and the fallback path all priced into the total.
Instrument | Evaluation and reliability calculatorsEvaluation Card Generator: Markdown and JSON on a Published Schema
Write one evaluation down so a second person can read the number correctly, or run it again. Fill the form, export a Markdown card and a JSON card that validates against a schema published here.
Instrument | Reliability testingGolden Set Version Tracker: Dataset Diff and Changelog
Diff two versions of your golden eval set in the browser, keep a dated changelog of what changed and why, and freeze the baseline metric values the next run is judged against.
Instrument | Reliability testingRelease Gate Designer: Go/No-Go Rules and Decision Log
Write the rule that turns your eval figures into PROMOTE, HOLD or ROLLBACK, get the verdict plus the rule that produced it, and keep a dated decision log and memo in your browser.
Instrument | Reliability testingChange Impact Worksheet: What to Rerun Before You Ship
Describe one change to a shipped AI system and get the rerun list it invalidates: which suites, slices and graders to run again, why each is on the list, and a change record for the ticket.
Instrument | Eval statisticsIndependent Two-Proportion Test Calculator for Eval Arms
Compare two pass rates measured on different case sets: a pooled two-proportion z test, Fisher exact when an arm is small, and an interval on the difference.