Topic
Reliability testing
Testing agents like software: harnesses, regression suites, and what a passing run does and does not prove.
Research
- Analysis
Tool output truncation: per-framework limits and what breaks
Tool output truncation cuts a tool's return value before the model reads it. A per-runtime table of the defaults, each read from its own docs or code, plus the downstream failures a silent cut causes.
- Analysis
Why AI agents hang: timeouts, stalls, and stops that don't
Agent runs hang in three shapes: an unbounded wait, an inactivity timeout a slow stream keeps alive, and a cancellation the work declines. Verified framework defaults, and how to tell them apart.
- Analysis
How coding agents fail: a trajectory study of CLI agents
Across 1,794 annotated CLI coding-agent trajectories, the decisive error landed at a median of step 7, locked in one step later, and stayed invisible until around step 16. What that means for evals.
- Analysis
Silent failures: when agents report success and are wrong
Four 2026 papers on silent failure in AI agents: how often failed runs carry explicit success claims, why LLM judges barely beat chance at catching them, and what detects false success instead.
- Analysis
Structured output failures across models and frameworks
Structured output breaks four ways: a schema rejected at compile, the framework swapping enforcement methods, generation truncating mid-object, and output that validates while carrying a wrong value.
- Study
Claude Fable 5 vs GPT-5.6 Sol vs Kimi K3 reliability benchmark
The full three-way benchmark behind our builder guide. Claude Fable 5, GPT-5.6 Sol and Kimi K3 on identical tasks, eight areas scored, every interval and caveat published.
- Analysis
The agent reliability testing checklist your green eval skips
A pre-flight agent reliability testing checklist: six checks with pass criteria and calculators, from failure-mode coverage and statistical power to pass^k, cascade budget, and judge calibration.
- Analysis
Make AI agents reliable in production by budgeting for failure
Reliable AI agents in production start with the compounding math: 0.95 per step over 20 steps is ~36% end to end. A topology-aware playbook of levers: retry, fallback, checkpoint, human-gate, degrade.
- Analysis
Agentic AI testing: the five dimensions one eval run cannot reach
How to design the test rather than pick a metric: run count and power, run-to-run variance, fault injection across topologies, and the protocol that turns five dimensions into a release decision.
- Analysis
How to measure agent reliability past a single pass rate
How to measure agent reliability with metrics that capture the consistency a single pass rate cannot: pass@k versus pass^k, a reliability@k suite aggregate, and a confidence interval on every rate.
- Analysis
AI agent reliability, from consistency to containment
AI agent reliability is a discipline of five properties: consistency, robustness, predictability, safety, and error propagation, with a map of where each is measured.
Guides
- Guide
Golden set blueprint: coding agents
Seven case families for coding agents: must-flip and must-not-break tests on separate denominators, solution leakage measured, and the test suite treated as the evaluator it is.
- Guide
Golden set blueprint: customer support
Seven case families, five slices, a coverage skeleton, and ten starter rows for testing a support agent, with a finance overlay and CSV/JSONL downloads.
- Guide
Golden set blueprint: regulated content
Seven case families for content under regulatory constraints: per-output constraint checks, risk-disclosure preservation, audience mismatch, and a pharma overlay. Not compliance advice.
- Guide
Golden set blueprint: research agents
Seven case families for research agents: citation existence, citation support, evidence strength, source coverage, and silent fabrication, each graded per claim against a named source.
- Guide
Golden set blueprint: transactional agents
Seven case families for agents that change world state: expected writes, confirmation integrity, silent tool failure, budget exhaustion, and a finance overlay.
- Guide
We scored Fable 5 vs Opus 5 vs Opus 4.8 and only the old one answered every call
While Claude Fable 5 and Claude Opus 5 are tied on our reliability benchmarks, Opus 4.8, surprisingly, is still the better choice for two specific kinds of work.
- Guide
We scored Fable 5 vs Sol vs Kimi K3 and the winner refused the most calls
Claude Fable 5 wins on points, GPT-5.6 Sol comes last and answers everything, and Kimi K3 is quick until it hangs. Pick by the failure your pipeline can absorb.
- Guide
Model fallback swaps the model unseen, and refusals log as success
When Claude Fable 5 refuses, it returns HTTP 200 with an empty content array, so error dashboards read it as a success, and with fallback on the model can swap to Opus 4.8 unseen. Instrument now.
- Guide
The OWASP AI agent security cheat sheet, read as a reliability spec
OWASP's AI Agent Security Cheat Sheet and the 2026 Agentic Top 10 name the same agent failures. Each maps to a reliability control you can test: least-privilege tools, isolated memory, bounded loops.
- Guide
AI observability proves the run finished. An eval proves it was right.
LangChain's State of Agent Engineering survey of 1,340 practitioners: 89% run observability, only 52.4% run offline evals. An offline eval is what scores whether the answer was right.
- Guide
How benchmarks get gamed, and how to check yours
BenchJack, a Berkeley auditing tool, found 219 flaws in ten popular agent benchmarks and gamed nine to near-perfect scores. Why a benchmark number is a claim about the harness, and how to check it.
Terms
-
Chaos engineering for AI agents
Chaos engineering for AI agents is the disciplined injection of controlled faults into an agent system to measure how far each fault propagates and what fraction of it the architecture contains, each result reported with a confidence interval.
-
Fault injection (agents)
Fault injection (agents) is a controlled experiment that introduces a chosen fault at a known point in an agent topology and measures how far it propagates, so its reach and containment become recorded quantities a system can be scored on.
-
reliability@k and pass^k
pass^k is the probability an agent solves all k runs of one task (closed form p^k). reliability@k is the lane's suite-level aggregate of pass^k: the mean across a representative task suite. It is the consistency counterpart to pass@k (best-of-k capability), not its inverse.
Calculators
- Calculator
Irreversible Action Inventory: What Your Agent Cannot Take Back
List every action your agent can take, mark what reverses it, and name the point of no return. A written verdict per row, never a score or a budget.
- Calculator
LLM Token Counter: Context Budget and Cost per Attempt
Count tokens for a named model, see what the request reserves in the context window once output is set aside, and price one attempt at your own rates. Exact where we bundle the encoding.
- Calculator
Human Review Threshold Optimizer for Confidence-Gated Automation
Derive the human review threshold from your own records: paste stated confidence and observed correctness, price one review against one missed failure, and read the cut that costs least.
- Calculator
AI SLO Error Budget Calculator: Split a Failure Budget Across Classes
Turn a completion-rate target and a task volume into a failure budget split by class, with a Wilson interval on each class observed rate against its allowance.
- Calculator
Prompt Robustness Analyzer: Perturbation Stability for LLM Evals
Is a prompt's pass-rate drop real or noise? Paste pass/fail results for a baseline prompt and its semantically equivalent variants and read each change with a paired interval and an exact McNemar p.
- Calculator
Golden Set Version Tracker: Dataset Diff and Changelog
Diff two versions of your golden eval set in the browser, keep a dated changelog of what changed and why, and freeze the baseline metric values the next run is judged against.
- Calculator
Release Gate Designer: Go/No-Go Rules and Decision Log
Write the rule that turns your eval figures into PROMOTE, HOLD or ROLLBACK, get the verdict plus the rule that produced it, and keep a dated decision log and memo in your browser.
- Calculator
Change Impact Worksheet: What to Rerun Before You Ship
Describe one change to a shipped AI system and get the rerun list it invalidates: which suites, slices and graders to run again, why each is on the list, and a change record for the ticket.
- Calculator
AI Model Deprecation Dates, Replacements and Breaking Changes
Retirement dates, vendor replacements and what breaks in your eval when you swap.
- Calculator
Agent Trace Analyzer: Step Counts, Repeated Calls, Handoff Edges
Upload an agent run trace and get step counts, repeated-call detection with index ranges, handoff edges, and per-class run tallies. Deterministic checks on the file you already have.
- Calculator
Context-of-Use Profiler: Turn a System Description into Evaluation Consequences
Describe what your AI system does, who it affects and how it fails, and get the failure families, case classes, slices and evidence expectations those answers demand.
- Calculator
LLM Latency Budget Calculator: TTFT and Token Time Across a Chain
Compose per-step time-to-first-token and token-stream latency into a chain deadline. See which step spends the budget, what kind of time it spends, and the timeout each layer should hold.
- Calculator
Evaluation Blueprint Builder: Write the Argument Before You Run the Eval
State each claim your evaluation will make, name the evidence that would support or refute it, and write the inference step connecting the two, all before any results exist.
- Calculator
Risk-to-Test Mapper: Which Failures Need Which Cases
Name the failures you care about, then see the test-case shapes, slices, evaluators and metrics each one needs. Every mapping says what it cannot establish. The mapping is a LatentEval synthesis.
- Calculator
Acceptance Criteria Builder: Set the Bar Before the Run
State each pass/fail criterion, the risk axes it covers, the reasoning behind the threshold and who set it. Get the sample size each one needs and export the rules for the release gate designer.
- Calculator
Acceptable Error Rate Calculator: The Reliability Bar Automation Needs
Invert your own manual cost, automation cost, and consequence cost into the minimum reliability that makes automating pay, and see if a measured rate clears it.
- Calculator
Eval Run Register: Preregistration and Run Cards for LLM Evals
Keep one row per eval run in this browser: freeze the plan before results exist, log what happened, and export a run card or a preregistration record for a PR.