Reference
Your agent did this: find the page that explains it
48 symptoms | 15 categories
Name the symptom in your own words and every result links to the calculator, research page or glossary entry that measures or defines it. The research hub sorts readers by three reliability properties it names for them; this page takes the phrase you already have and returns the page.
Loops and non-termination
3 symptoms- Agent stuck in a loop Calculator: set an iteration cap and a dollar ceiling from your observed step counts
- Agent repeats the same tool call Trace check: upload a trace and see repeated-action sequences flagged
- Loop detection for an agent Calculator: cap planner returns the config snippet that stops a loop before the budget runs out
Runaway costs
3 symptoms- Runaway agent costs Calculator: price one acceptable result after retries, failed attempts and fallback
- Unexpected LLM bill Calculator: see the per-attempt token cost against the budget for the whole task
- Agent cost keeps climbing Calculator: set an iteration cap and a dollar ceiling so costs stop before the task does
Silent tool failures
3 symptoms- Agent tool call fails silently Research: how false-success failures hide inside passing eval runs
- Agent calls the wrong tool Auditor: paste a tool surface and see which names overlap or which schema is ambiguous
- Tool error the agent ignored Glossary: what task-verification failure means and how to test for it
Nondeterminism and flaky results
3 symptomsSilent model changes
3 symptoms- Silent model regression Guide: how to log and detect model-level changes before they reach your eval
- How to prove a model changed Calculator: paired-run comparison with bootstrap intervals on the difference
- Score dropped after a model update Calculator: McNemar test on the same items scored before and after the change
Handoff and context loss
3 symptoms- Context loss between agents Glossary: what context-handoff loss means and the conditions that cause it
- Agent-to-agent handoff failure Trace check: upload a trace and inspect the handoff step for dropped fields
- Information lost in a handoff Glossary: what information withholding between agents looks like and how to detect it
Hangs, timeouts and unfinished runs
4 symptoms- Agent times out Research: why agents hang, which orchestrators carry which default timeouts, and what to set
- Agent never finishes the task Glossary: what a timeout budget is and how to derive one from a chain deadline
- Agent stopped before the task was done Glossary: what premature termination means and when an iteration cap causes it
- Steps take too long and the chain misses its deadline Calculator: compose per-step TTFT and token latency into a chain deadline and see which step spends the budget
Retry storms and rate limits
3 symptoms- Agent retry storm Glossary: what a retry storm is and how agents trigger one
- Rate limit hit after retries Calculator: compare retry, backoff, and fallback strategies on delivered correctness and cost
- Exponential backoff for an LLM agent Glossary: what exponential backoff with jitter means and the standard parameters
Tool-surface and MCP problems
4 symptoms- MCP server not working Auditor: paste one manifest and see tool-count, schema-size and overlap findings
- Too many MCP tools Auditor: the tool-count and schema-token readout tells you which tools to cut
- MCP tool schema breaking change Auditor: paste two manifests and see every field that moved, was added, or was removed
- Agent picks the wrong tool from too many Guide: what an MCP gateway does and when aggregation becomes the failure mode
Deprecation and migration
3 symptoms- Model deprecation date and migration path Reference: retirement dates, replacements and what breaks, maintained and dated
- Eval harness is shutting down Migration kit: export your suite and check whether scores moved because the system changed or the harness did
- Score changed after switching harness Migration kit: paired before-and-after comparison on the same items, inside the overlap window
Judge disagreement
4 symptoms- Judge disagrees with human labels Router: upload scores and human labels, see which bias mode explains the gap, and follow the link to the calculator that measures it
- Judge grades its own output higher Glossary: what self-preference bias means and how to measure it
- Judge prefers longer answers Glossary: what verbosity bias means and how to separate length from quality
- Judge and humans agree less than expected Calculator: chance-corrected agreement between two raters with an interval and a band
Cascade and error propagation
4 symptoms- One wrong answer becomes three Research: how a single fault propagates through a multi-agent pipeline and what amplifies it
- Production agent failures Research: what breaks in production agents and where each failure class is measured
- Cascade drops end-to-end reliability Calculator: end-to-end success rate from per-step reliability, and the per-step budget your target demands
- Simulator: price a chain when its steps share a cause, with the retry ceiling shared failures impose
RAG and retrieval failures
3 symptoms- Agent answers from the wrong document Research: where RAG pipelines fail, from retrieval to generation
- RAG answers are not grounded Glossary: what groundedness means and how to measure it in a RAG pipeline
- Retrieval quality is unknown Calculator: precision at k, recall at k, nDCG, MRR and MAP from your retrieved and relevant sets
Structured-output failures
2 symptomsScore moved: is it real?
3 symptoms- Eval score moved: is it real? Research: statistical tests for eval differences, when each applies and what sample size they need
- Need a confidence interval on a pass rate Calculator: Wilson and Clopper-Pearson intervals on a pass rate from your counts
- Which test compares two eval runs Calculator: McNemar for paired items, two-proportion for independent arms