LatentEval

INSTRUMENT | Reliability testing

LLM Latency Budget Calculator: TTFT and Token Time Across a Chain

1 cited source

Compose per-step time-to-first-token and token-stream latency into a chain deadline. See which step spends the budget, what kind of time it spends, and the timeout each layer should hold.

Enter each step's time-to-first-token, output token count and inter-token latency. The calculator composes them into a chain total against an optional deadline, names the step that spends the budget, and derives the timeout each layer should hold from the framework defaults in our agent hangs and timeout research. The error-budget calculator spends a failure budget, counting how many tasks a target lets you lose; this page spends a time budget, counting how much of a deadline each step consumes.

Deadline and percentile

The whole run's wall-clock ceiling. Leave blank to see the chain total without a fits/misses verdict.

Check this value.

Which percentile are these numbers? The calculator does not guess. It carries your label through every result.

Check this value.

Steps
StepTTFT (s)Output tokensInter-token latency (ms/tok)Tool/fixed time (s)AttemptsBackoff (s) Actions
  • Step name. A short label. Leave blank for a numbered position.
  • TTFT (s). Time to first token, including queueing and prefill. Measure on your own traffic.
  • Output tokens. Every token the model generates, including reasoning tokens you never see. A model that spends 19,000 tokens thinking before answering uses 19,000 here, not the 64 visible ones.
  • Inter-token latency (ms/tok). The gap between output tokens during decoding. Also called TPOT (time per output token).
  • Tool/fixed time (s). Non-model time in this step: the tool call, a network hop, orchestration overhead.
  • Attempts. Planned attempts including the first. 1 means no retry. The total counts every attempt.
  • Backoff (s). Wait time between attempts. Applied between attempts, not after the last one.

Chain total, with retries

30.8 s

misses | Headroom -755 ms | Utilization 102.5% | Percentile label p95

misses

The chain totals 30.8 s and misses the 30.0 s deadline by 755 ms; the largest share, 68.9%, is spent by draft the answer, and it first runs past the deadline at check the answer.

Chain total, single attempt

26.5 s

Headroom

-755 ms

Per step | p95

StepModel timeAttempt timeStep timeSingle attemptShare of chainTTFT shareStream shareFixed shareSpend kindRemaining at entry (timeout)Solo allowanceOverrunMax tokens
route the request 1.2 s 1.2 s 1.2 s 1.2 s 3.8% 33.9% 66.1% 0.0% stream-dominant 30.0 s 425 ms 28.8 s 2
retrieve context 0 ms 350 ms 350 ms 350 ms 1.1% 0.0% 0.0% 100.0% fixed-dominant 28.8 s -405 ms 28.5 s -
draft the answer (spends most) 21.2 s 21.2 s 21.2 s 21.2 s 68.9% 5.7% 94.3% 0.0% stream-dominant 28.5 s 20.4 s 7.3 s 769
check the answer (first overrun) 3.8 s 3.8 s 8.1 s 3.8 s 26.2% 21.2% 78.8% 0.0% stream-dominant 7.3 s 7.3 s -755 ms -
  • retrieve context: Negative: the other steps alone exceed the deadline.
  • retrieve context: Not available at zero inter-token latency.
  • check the answer: Retry overhead exceeds the attempt time.
  • check the answer: Not available when the step has retries.

These figures are a four-step illustration. Replace every cell with your own numbers.

Export

stepMs = attempts x (ttft + max(tokens - 1, 0) x itl + fixed) + (attempts - 1) x backoffHow?

How this is calculated

The chain total is a sum of per-step percentile assumptions. It is not the chain's own percentile at that level and it is not an upper bound on it.

Output tokens are counted with a minus-one convention: the first token is already paid for by TTFT, so only the remaining tokens each cost one inter-token gap. At one output token the step time equals TTFT exactly. At 800 tokens the difference from the loose convention is one inter-token gap, about 0.1%; at one token the loose convention double-counts the entire gap.

The table shows two allowance columns when a deadline is set. Remaining at entry is the time left when the chain actually reaches a step, subtracting only the work that ran before it; this is the timeout to hand that layer. Solo allowance is the room the step would have if every other step ran exactly as planned; this is the input to the token ceiling. The two numbers coincide only for the last step.

Hidden reasoning tokens count as output tokens. A model that spends 19,000 tokens thinking before producing 64 visible tokens uses 19,000 in this calculator, because every generated token occupies one inter-token gap regardless of whether the caller sees it. The gap between the visible token count and the billed token count is the measurement a timeout budget exists to surface.

Formula: stepMs = attempts x (ttft + max(tokens - 1, 0) x itl + fixed) + (attempts - 1) x backoff

Questions

Where do these default numbers come from?

The form opens on a four-step worked example: a router, a retrieval call, a draft step and a check step with one retry. The numbers are round figures chosen to show every column, including the retry overhead and the first-overrun marker. They are an illustration, not a measurement of any system.

Why does the token ceiling say 'n/a' for my step?

Three conditions suppress the token ceiling. A step with retries has no single attempt to ceiling against. A step at zero inter-token latency produces tokens instantly, so no token count is limiting. And a step whose solo allowance is smaller than its TTFT plus its fixed time has no room for even one token.

How does the retry row relate to the single-attempt row?

The step time includes every attempt and the backoff between them. The single-attempt time is what one try costs before any retry. The difference is the retry overhead, and the spend-kind split describes one attempt so it stays comparable across steps with different retry counts.

Sources

  1. AgentServeSim: hardware-aware simulator for multi-turn LLM agent serving (Rajib, Zheng, Lou)arXiv Retrieved