LatentEval

INSTRUMENT | Reliability testing

Agent Trace Analyzer: Step Counts, Repeated Calls, Handoff Edges

Upload an agent run trace and get step counts, repeated-call detection with index ranges, handoff edges, and per-class run tallies. Deterministic checks on the file you already have.

Upload a trace file from a finished agent run. The page counts steps, tool calls, repeated-call runs with index ranges, and handoff edges, built on the step-level failure patterns documented in the trajectory failure analysis. It reads OpenTelemetry GenAI spans, the OpenAI Agents SDK export, and plain tool-use message logs.

The run-class counts link to the pass-rate interval calculator for the rate and its interval. To record a run rather than read one, use the eval run register. To set an iteration cap from many runs' step counts, use the agent loop cap calculator.

Analysis mode

No trace content leaves your browser. The file is read and counted on this page, and no part of it is sent anywhere. Limits: 10 MB and 200,000 lines per file. Over either one the page says so and reads nothing.

JSON, JSONL, or NDJSON. Drag one here, or use the box below.

Check this value.

Detected from the file. Override if the detection is wrong.

Check this value.

How many consecutive identical calls count as a repeated run. Your choice, not a standard.

Check this value.

Longest repeating pattern to search for, in steps.

Check this value.

How many consecutive steps using only previously seen signatures count as churning. Your choice, not a standard.

Check this value.

Upload or paste a trace file. The page counts steps, tool calls, repeated-call runs with index ranges, handoff edges, and per-class run tallies. It reports what the file records and nothing else.

How this is calculated

Steps are ordered by timestamp when every step has one. When some timestamps are missing, or the format provides none, steps stay in file order for the entire run. The rule used is shown in the header.

A step's signature is the combination of its kind, its tool name and its canonical argument serialization. Two calls to the same tool with the same arguments produce the same signature. Three checks run over the signature sequence of each run:

  • Immediate repeat (L1). The longest consecutive run of one signature. Flags when the run length reaches the repeat threshold you set.
  • Cycle repeat (L2). The longest consecutive repetition of a multi-step pattern. For each window length from 2 up to the max cycle length, the page finds the maximum number of consecutive repetitions. Windows whose steps are all the same signature are skipped so L1 runs are not double-reported. Flags when a pattern repeats at least twice after its first appearance.
  • No-new-signature window (L3). The longest run of consecutive steps that use only signatures the run has already seen. A new tool call resets the window. Flags when the window length reaches the threshold you set.

Each run is classified into exactly one class. The first rule that matches wins:

ClassRule
UnknownThe format carried no status signal anywhere in the run: no step has a recorded status, the root span has no status, and the last step has no output.
ErroredAny step has a recorded error status, or the root span status is error.
No outputNot errored, and the last step produced no output.
CompletedA status signal exists, nothing errored, and the last step produced output.

One trace is a sample of one. To put an interval on a class rate measured across runs, use the repeated-run variance planner to decide how many runs the interval needs.

Questions

Which trace formats are accepted?

OpenTelemetry GenAI spans exported as OTLP/JSON, the OpenAI Agents SDK span export, and plain tool-use message JSONL. The format is detected from the file. If detection is wrong, override it with the format selector.

What does "not applicable" mean for handoffs?

The plain message JSONL format does not carry handoff information. When the file uses that format, handoff counts are reported as not applicable rather than as zero, because zero would imply the page looked and found none.

Why does the page not say what caused a failure?

The page counts what the trace records. It does not interpret, diagnose or attribute a cause. A step that errored is counted as errored. What went wrong is a question for the trace itself, not for a counter reading it.

What is the handoff contract field for?

It lists what the receiver of a handoff was supposed to get. Each line is a literal string or a regex pattern. The page checks each handoff's inbound payload against every item in the contract and reports which items were missing. Without a contract, the page still counts handoff edges, but the loss measurement has nothing to measure against.