INSTRUMENT | Propagation & containment
Agent Loop Cap Calculator
1 cited source
Paste your agent's step counts. Get the iteration cap your coverage target implies, the share of work the cap would cut, and a dollar ceiling for your unattended window.
A 95th-percentile cap of 18 steps from twenty completed runs truncates 5% of healthy work, with an order-statistic interval of 1.2% to 24.9%. This page derives that cap from your own step counts, prices the ceiling for an unattended window, and emits the config snippet. It rests on the framework bounds documented in our agent hangs and timeout research. For retries within a single step, use the retry strategy optimizer. To read one trace file and see step counts and repeated calls before aggregating, use the agent trace analyzer.
Showing your last valid result. Update the inputs above to recompute.
Iteration cap
18
95% coverage over 20 completed runs.
One completed run sets the same cap at every coverage target.
- Completed runs in the sample
- 20
- Expected steps per completed run, under the cap
- 8.85
- Observed truncation rate
- 5.0% (1/20)
- Order-statistic 95% interval
- 1.2% to 24.9%
- Budget-derived cap
- –
Candidate caps
| Coverage | Cap | Truncation | Interval | Per-run ceiling | Window ceiling |
|---|---|---|---|---|---|
| 50% | 7 | 50.0% (10/20) | 31.5% to 72.8% order statistic 95% | – | – |
| 90% | 14 | 10.0% (2/20) | 3.2% to 31.7% order statistic 95% | – | – |
| 95% (yours) | 18 | 5.0% (1/20) | 1.2% to 24.9% order statistic 95% | – | – |
| 99% | 31 | 0.0% (0/20) | 0.1% to 16.8% order statistic 95% | – | – |
Parameter names below were read from each framework's documentation. Defaults move; the dated table is on the agent hangs and timeout research page.
# agent loop cap policy max_steps_per_run: 18 coverage_target_pct: 95 sample_runs: 20 # framework parameter names for the step cap; no defaults are quoted here # langgraph: recursion_limit # crewai: max_iter # openai_agents: max_turns
{
"max_steps_per_run": 18,
"coverage_target_pct": 95,
"sample_runs": 20,
"framework_parameters": {
"langgraph": "recursion_limit",
"crewai": "max_iter",
"openai_agents": "max_turns"
}
} k = ceil(q * n); cap = x(k)How?
How this is calculated
The cap. The counts are sorted ascending as x(1) to x(n), the rank k is ceil(q * n) clamped to [1, n], and the cap is x(k). This is the nearest-rank empirical quantile, Hyndman and Fan type 1. Three properties make nearest-rank the right definition here: the cap is always an integer and always a value someone observed, which a config parameter requires; it is the smallest integer c with the share of runs at or under c reaching q, so the coverage promise is exact; and it is conservative versus the interpolating definitions most libraries default to, which silently truncate more runs than the reader asked for.
The resolution limit. With n runs the highest coverage the sample can resolve is 1 minus 1/n. When ceil(q * n) equals n, the cap is the observed maximum, and the page says so beside the number.
Two intervals, and they are not interchangeable. A cap derived from coverage is picked out of the sample, so the truncated count is n minus k by construction whenever x(k) is unique. A binomial interval on a count the selection rule already fixed reports sampling error the sample cannot have. The quantity you want is the share of future healthy runs this cap will cut, and for the k-th order statistic that has an exact distribution-free form, Beta(n - k + 1, k). A cap fixed without looking at the sample, which is what the budget-derived cap is, is a genuine binomial count against a fixed threshold, and that row gets a Wilson interval instead. The order-statistic form is exact for a continuous distribution and conservative when values tie at the cap, which integer step counts make common, so its upper bound is never too small. At k equal to n it becomes Beta(1, n), which is why a cap set at the observed maximum still carries a non-zero upper bound instead of a false zero.
Spend. The per-run ceiling is cap times the cost per step, and the window ceiling is that times the runs started in the window. Both expected figures are the capped mean over completed runs and the page labels them that way. The sample holds no run that hung; a hung run spends the full cap; so both are a lower bound on real spend and the gap widens with the runaway rate. Read the pair: true window spend sits between the window expected cost and the window ceiling, and only the ceiling is a bound.
The inverse direction. A window budget B becomes floor(B / (runs in the window times the cost per step)), and the truncation rate is then re-run at that cap. When the result falls below 1 the budget cannot fund one step of every run, and the page says that rather than returning a zero.
Censoring. The step counts come from runs that finished. A run that looped forever contributed nothing to the sample. The cap marks the point past which a run stops looking like normal work. It is not a prediction of how far a broken loop would have gone. That is why the window ceiling matters more than the cap: the cap is a shape statement about healthy runs, the ceiling is the bound on the bill.
The worked example. Twenty completed runs at a 95% coverage target give k = 19 and a cap of 18 steps. One run in twenty is truncated, a rate of 5.0%, with an order-statistic 95% interval of 1.2% to 24.9%. At $0.042 a step the per-run ceiling is $0.756 and 720 runs put a $544.32 ceiling on the window. Capping moves the expected cost of a completed run from $0.399 to $0.3717, about 6.8%, so the ceiling rather than the saving is the number worth carrying away.
A cap bounds the bill. It says nothing about what the run can do before it stops. Before you raise one, check which of the run's steps can do something you cannot take back with the agent action risk matrix, and read the mechanism the cap actually triggers under premature termination.
Formula: k = ceil(q * n); cap = x(k)
Questions
Why nearest-rank, and not the percentile my library returns?
The cap is always an integer and always a value someone observed, which a config parameter requires. It is the smallest integer c with the share of runs at or under c reaching your coverage target, so the coverage promise is exact. And it is conservative versus the interpolating definitions most libraries default to, which silently truncate more runs than you asked for. Two tools handed the same data and asked for a p95 return different values, which is why the definition is named on this page.
Why does the budget row use a different interval?
A cap derived from coverage is picked out of your sample, so the count above it is fixed by the selection rule and a binomial interval on it would report sampling error the sample cannot have. A cap that came from a budget was fixed without looking at the sample, so the count above it is a genuine binomial count against a fixed threshold. Wilson is correct there and nowhere else on this page.
Does the cap stop a runaway agent?
The step counts come from runs that finished. A run that looped forever contributed nothing to the sample. The cap marks the point past which a run stops looking like normal work. It is not a prediction of how far a broken loop would have gone. That is why the window ceiling matters more than the cap: the cap is a shape statement about healthy runs, the ceiling is the bound on the bill.
Why does the snippet name parameters but no default values?
Defaults move, and a page that quotes them becomes a table someone has to maintain. The parameter names are stable; the values are not. The dated table lives on the agent hangs and timeout research page, which is the page that carries the reading date.
What if I only have a handful of finished runs?
With n runs the highest coverage the sample can resolve is 1 minus 1/n. Ask for more than that and the cap is simply the observed maximum, which bounds nothing a healthy run has not already done. With a single run the cap is that value at every coverage target, and the interval runs from 2.5% to 97.5%, which is the honest reading of one run.