LatentEval

INSTRUMENT | eval

Evaluation Card Generator: Markdown and JSON on a Published Schema

4 cited sources

Write one evaluation down so a second person can read the number correctly, or run it again. Fill the form, export a Markdown card and a JSON card that validates against a schema published here.

An evaluation card is what a second person needs to read your number correctly, or run it again. 96.5 percent of published results miss a basic reproducibility field. It builds on our house methodology and our research into LLM evals.

Fill the form and export Markdown and JSON that validate against a schema published here. To check a card someone else wrote, use the reproducibility checklist. For a running record across many runs, use the run register.

This form opens on a finished card: our own pre-registered routing study. Every value in it is ours, not yours.

Subject

Check this value.

Check this value.

The pinned snapshot. Where the provider exposes none, say how you reached it and when.

Check this value.

Check this value.

Check this value.

Check this value.

Check this value.

Check this value.

Task

The one sentence this number is evidence for.

Check this value.

Check this value.

What the metric stands in for.

Check this value.

Check this value.

Check this value.

Check this value.

Data

Check this value.

Check this value.

Check this value.

Check this value.

The digest of the bytes that were scored: 64 lowercase hex characters. A named file with no digest can change under the number.

Check this value.

Check this value.

Check this value.

Check this value.

Check this value.

Check this value.

Check this value.

Check this value.

Check this value.

Check this value.

What you did to check the subject had not already seen these items.

Check this value.

Check this value.

Method

Check this value.

Check this value.

Which way is better. A metric with no direction cannot be compared.

Check this value.

Check this value.

Check this value.

Check this value.

How a raw output became a score.

Check this value.

Check this value.

Check this value.

Check this value.

Check this value.

Check this value.

Check this value.

Check this value.

Zero is a real seed. Leave it blank only if none was set.

Check this value.

Check this value.

Other settings

SettingValue Actions

Check this value.

Check this value.

Check this value.

The step, tool-call, wall-clock or token budget the run was held to.

Check this value.

Check this value.

Judge

Both the name and the version, or neither. A judge-free run leaves this empty rather than filling a placeholder.

Check this value.

Check this value.

Check this value.

Check this value.

Check this value.

Harness

Check this value.

Check this value.

Check this value.

Result

The number and what it means, in one line.

Check this value.

Check this value.

Figures

FigureValueIntervalInterval kindn Actions

Check this value.

Limits

Limits

Where this evaluation does not reach Actions

Check this value.

Provenance

First-party and third-party reports fill these fields at very different rates.

Check this value.

Check this value.

Check this value.

Check this value.

Check this value.

Check this value.

Cost
Cost

What the run cost you. Zero is a real answer.

Check this value.

Check this value.

Check this value.

Check this value.

of 35 fields recorded

27

Core complete

Two counts and the gaps. Counts, never a score: a completeness percentage is a quality rating in disguise.

What the card recordsCount
Fields recorded 27 of 35
Core fields recorded 15 of 15
Not recorded data.license: The scored task set is not published, so it carries no license of its own. The results file is CC BY 4.0. · data.contaminationCheck: No contamination check was run. The tasks were written for this study rather than drawn from a public benchmark. · method.seed: No seed was set or recorded. The write-up notes that thinking is non-deterministic regardless. · method.evalPlan: There was no agent scaffold. Each task was a single call. · method.evalLimits: No step, tool-call or wall-clock budget applied to a single-turn call with no tools. · method.judge: Scoring was deterministic: exact match, unit test and tool-argument match, with no model judge. · method.harness: The study was pre-registered, run and scored by hand rather than through a named harness. · cost: Token spend for the sweep was never itemized, so no figure is recorded here.
Markdown card
# Evaluation card

## Subject

- **Name:** Fable 5
- **Version:** As served by the API on 2026-07-07
- **Provider:** Anthropic
- **Access method:** api

## Task

- **Claim:** No measurable capability separation on short checkable tasks.
- **Construct:** Capability on short, deterministic, machine-checkable tasks.
- **Intended use:** Routing short checkable work between model tiers.

## Data

- **Source:** Model routing and the refusal tax: results dataset, https://latenteval.ai/benchmarks/model-routing-refusal-tax-results.csv
- **Version:** v3.withRescue
- **SHA-256:** d90ed5705cb0799051110f5fa4213828c4fc1901f591753b233af0f525f693ae
- **Size:** 28
- **Split:** low effort, rescue on
- **As of:** 2026-07-07
- **License:** Not recorded: The scored task set is not published, so it carries no license of its own. The results file is CC BY 4.0.
- **Contamination check:** Not recorded: No contamination check was run. The tasks were written for this study rather than drawn from a public benchmark.

## Method

- **Metric:** pass rate
- **Direction:** higher-is-better
- **Scale:** 0 to 1
- **Parser:** Deterministic exact match, unit test, or tool-argument match. No model judge.
- **Runs:** 1
- **Temperature:** 1
- **Max tokens:** 8192
- **Seed:** Not recorded: No seed was set or recorded. The write-up notes that thinking is non-deterministic regardless.
- **Other settings:** effort = low; server-side fallback = on, Fable 5 to Opus 4.8
- **Eval plan:** Not recorded: There was no agent scaffold. Each task was a single call.
- **Eval limits:** Not recorded: No step, tool-call or wall-clock budget applied to a single-turn call with no tools.
- **Judge:** Not recorded: Scoring was deterministic: exact match, unit test and tool-argument match, with no model judge.
- **Harness:** Not recorded: The study was pre-registered, run and scored by hand rather than through a named harness.

## Result

- **Headline:** 96.4 percent with rescue on, 95 percent Wilson interval 82.3 to 99.4.
- **Figures:** Pass rate, rescue on: 0.9643, interval 0.8229 to 0.9937 (ci), n = 28; Rescue fired: 20 of 28 calls, n = 28; Refusal rate, low effort, rescue off: 21 of 28 calls, n = 28

## Limits

- **Limits:** Short, deterministic, machine-checkable correctness only. Long-horizon agent work is unmeasured here.; Small n: 14 pruned tasks, two runs each, 28 trials per arm. The intervals are wide and directional only.; One provider, one task family, one 2026-07-02 classifier snapshot. The run config is not published, so the temperature and max-token values on this card cannot be recomputed from the results file.

## Provenance

- **Reporter:** first-party
- **Run by:** latenteval
- **Run on:** 2026-07-07

## Cost

- **Cost:** Not recorded: Token spend for the sweep was never itemized, so no figure is recorded here.

## Schema

- **Schema version:** v1
- **Conforms to:** https://latenteval.ai/schemas/evaluation-card/v1.json
- **Aligned with:** Evaluation Cards, the four signals and the minimal reproducibility fields, https://arxiv.org/html/2606.09809v1, read 2026-08-27
Export the card
JSON card
{
  "docType": "evaluation-card",
  "schemaVersion": "1.0",
  "title": "Evaluation card",
  "generatedAt": "2026-08-31",
  "sections": [
    {
      "heading": "Subject",
      "fields": [
        {
          "label": "Name",
          "value": "Fable 5"
        },
        {
          "label": "Version",
          "value": "As served by the API on 2026-07-07"
        },
        {
          "label": "Provider",
          "value": "Anthropic"
        },
        {
          "label": "Access method",
          "value": "api"
        }
      ]
    },
    {
      "heading": "Task",
      "fields": [
        {
          "label": "Claim",
          "value": "No measurable capability separation on short checkable tasks."
        },
        {
          "label": "Construct",
          "value": "Capability on short, deterministic, machine-checkable tasks."
        },
        {
          "label": "Intended use",
          "value": "Routing short checkable work between model tiers."
        }
      ]
    },
    {
      "heading": "Data",
      "fields": [
        {
          "label": "Source",
          "value": "Model routing and the refusal tax: results dataset, https://latenteval.ai/benchmarks/model-routing-refusal-tax-results.csv"
        },
        {
          "label": "Version",
          "value": "v3.withRescue"
        },
        {
          "label": "SHA-256",
          "value": "d90ed5705cb0799051110f5fa4213828c4fc1901f591753b233af0f525f693ae"
        },
        {
          "label": "Size",
          "value": 28
        },
        {
          "label": "Split",
          "value": "low effort, rescue on"
        },
        {
          "label": "As of",
          "value": "2026-07-07"
        },
        {
          "label": "License",
          "value": "Not recorded: The scored task set is not published, so it carries no license of its own. The results file is CC BY 4.0."
        },
        {
          "label": "Contamination check",
          "value": "Not recorded: No contamination check was run. The tasks were written for this study rather than drawn from a public benchmark."
        }
      ]
    },
    {
      "heading": "Method",
      "fields": [
        {
          "label": "Metric",
          "value": "pass rate"
        },
        {
          "label": "Direction",
          "value": "higher-is-better"
        },
        {
          "label": "Scale",
          "value": "0 to 1"
        },
        {
          "label": "Parser",
          "value": "Deterministic exact match, unit test, or tool-argument match. No model judge."
        },
        {
          "label": "Runs",
          "value": 1
        },
        {
          "label": "Temperature",
          "value": 1
        },
        {
          "label": "Max tokens",
          "value": 8192
        },
        {
          "label": "Seed",
          "value": "Not recorded: No seed was set or recorded. The write-up notes that thinking is non-deterministic regardless."
        },
        {
          "label": "Other settings",
          "value": [
            "effort = low",
            "server-side fallback = on, Fable 5 to Opus 4.8"
          ]
        },
        {
          "label": "Eval plan",
          "value": "Not recorded: There was no agent scaffold. Each task was a single call."
        },
        {
          "label": "Eval limits",
          "value": "Not recorded: No step, tool-call or wall-clock budget applied to a single-turn call with no tools."
        },
        {
          "label": "Judge",
          "value": "Not recorded: Scoring was deterministic: exact match, unit test and tool-argument match, with no model judge."
        },
        {
          "label": "Harness",
          "value": "Not recorded: The study was pre-registered, run and scored by hand rather than through a named harness."
        }
      ]
    },
    {
      "heading": "Result",
      "fields": [
        {
          "label": "Headline",
          "value": "96.4 percent with rescue on, 95 percent Wilson interval 82.3 to 99.4."
        },
        {
          "label": "Figures",
          "value": [
            "Pass rate, rescue on: 0.9643, interval 0.8229 to 0.9937 (ci), n = 28",
            "Rescue fired: 20 of 28 calls, n = 28",
            "Refusal rate, low effort, rescue off: 21 of 28 calls, n = 28"
          ]
        }
      ]
    },
    {
      "heading": "Limits",
      "fields": [
        {
          "label": "Limits",
          "value": [
            "Short, deterministic, machine-checkable correctness only. Long-horizon agent work is unmeasured here.",
            "Small n: 14 pruned tasks, two runs each, 28 trials per arm. The intervals are wide and directional only.",
            "One provider, one task family, one 2026-07-02 classifier snapshot. The run config is not published, so the temperature and max-token values on this card cannot be recomputed from the results file."
          ]
        }
      ]
    },
    {
      "heading": "Provenance",
      "fields": [
        {
          "label": "Reporter",
          "value": "first-party"
        },
        {
          "label": "Run by",
          "value": "latenteval"
        },
        {
          "label": "Run on",
          "value": "2026-07-07"
        }
      ]
    },
    {
      "heading": "Cost",
      "fields": [
        {
          "label": "Cost",
          "value": "Not recorded: Token spend for the sweep was never itemized, so no figure is recorded here."
        }
      ]
    },
    {
      "heading": "Schema",
      "fields": [
        {
          "label": "Schema version",
          "value": "v1"
        },
        {
          "label": "Conforms to",
          "value": "https://latenteval.ai/schemas/evaluation-card/v1.json"
        },
        {
          "label": "Aligned with",
          "value": "Evaluation Cards, the four signals and the minimal reproducibility fields"
        },
        {
          "label": "Alignment read on",
          "value": "2026-08-27"
        }
      ]
    }
  ],
  "data": {
    "card": {
      "schemaVersion": "v1",
      "conformsTo": "https://latenteval.ai/schemas/evaluation-card/v1.json",
      "externalAlignment": {
        "vocabulary": "Evaluation Cards, the four signals and the minimal reproducibility fields",
        "url": "https://arxiv.org/html/2606.09809v1",
        "checkedOn": "2026-08-27"
      },
      "subject": {
        "name": "Fable 5",
        "version": "As served by the API on 2026-07-07",
        "provider": "Anthropic",
        "accessMethod": "api"
      },
      "task": {
        "claim": "No measurable capability separation on short checkable tasks.",
        "construct": "Capability on short, deterministic, machine-checkable tasks.",
        "intendedUse": "Routing short checkable work between model tiers."
      },
      "data": {
        "source": "Model routing and the refusal tax: results dataset, https://latenteval.ai/benchmarks/model-routing-refusal-tax-results.csv",
        "version": "v3.withRescue",
        "sha256": "d90ed5705cb0799051110f5fa4213828c4fc1901f591753b233af0f525f693ae",
        "size": 28,
        "split": "low effort, rescue on",
        "asOf": "2026-07-07",
        "license": null,
        "contaminationCheck": null
      },
      "method": {
        "metric": "pass rate",
        "direction": "higher-is-better",
        "scale": "0 to 1",
        "parser": "Deterministic exact match, unit test, or tool-argument match. No model judge.",
        "runs": 1,
        "temperature": 1,
        "maxTokens": 8192,
        "seed": null,
        "otherParams": [
          {
            "key": "effort",
            "value": "low"
          },
          {
            "key": "server-side fallback",
            "value": "on, Fable 5 to Opus 4.8"
          }
        ],
        "evalPlan": null,
        "evalLimits": null,
        "judge": null,
        "harness": null
      },
      "result": {
        "headline": "96.4 percent with rescue on, 95 percent Wilson interval 82.3 to 99.4.",
        "figures": [
          {
            "label": "Pass rate, rescue on",
            "value": "0.9643",
            "interval": "0.8229 to 0.9937",
            "intervalKind": "ci",
            "n": 28
          },
          {
            "label": "Rescue fired",
            "value": "20 of 28 calls",
            "interval": null,
            "intervalKind": null,
            "n": 28
          },
          {
            "label": "Refusal rate, low effort, rescue off",
            "value": "21 of 28 calls",
            "interval": null,
            "intervalKind": null,
            "n": 28
          }
        ]
      },
      "limits": [
        "Short, deterministic, machine-checkable correctness only. Long-horizon agent work is unmeasured here.",
        "Small n: 14 pruned tasks, two runs each, 28 trials per arm. The intervals are wide and directional only.",
        "One provider, one task family, one 2026-07-02 classifier snapshot. The run config is not published, so the temperature and max-token values on this card cannot be recomputed from the results file."
      ],
      "provenance": {
        "reporter": "first-party",
        "runBy": "latenteval",
        "runOn": "2026-07-07"
      },
      "cost": null,
      "unrecorded": [
        {
          "path": "data.license",
          "reason": "The scored task set is not published, so it carries no license of its own. The results file is CC BY 4.0."
        },
        {
          "path": "data.contaminationCheck",
          "reason": "No contamination check was run. The tasks were written for this study rather than drawn from a public benchmark."
        },
        {
          "path": "method.seed",
          "reason": "No seed was set or recorded. The write-up notes that thinking is non-deterministic regardless."
        },
        {
          "path": "method.evalPlan",
          "reason": "There was no agent scaffold. Each task was a single call."
        },
        {
          "path": "method.evalLimits",
          "reason": "No step, tool-call or wall-clock budget applied to a single-turn call with no tools."
        },
        {
          "path": "method.judge",
          "reason": "Scoring was deterministic: exact match, unit test and tool-argument match, with no model judge."
        },
        {
          "path": "method.harness",
          "reason": "The study was pre-registered, run and scored by hand rather than through a named harness."
        },
        {
          "path": "cost",
          "reason": "Token spend for the sweep was never itemized, so no figure is recorded here."
        }
      ]
    }
  },
  "provenance": "This card is still the page's worked example, our own routing study, rather than a card someone filled in."
}
Export the card
Load a card back in

Paste the JSON export. The form returns to the state it was in, reasons included. A file from a different document type is refused by name.

recorded = fields with a value · core recorded = core fields with a valueHow?

How this is calculated

The two counts. There are 35 fields on a card and 15 of them are core. Both totals are read off the field list itself, so neither can drift from the fields you can see. A field counts as recorded when it holds a value. Zero is a value: a temperature of 0 and a seed of 0 are recorded, not empty, and that distinction is the single most common way a card like this goes wrong.

What the counts do not mean. They are counts of fields, not a rating. A card with 30 of 35 is not better than one with 25 in any sense this page will assert. There is no percentage here, no grade and no badge, on purpose: a completeness score invites a reader to compare two evaluations on how much paperwork came with them.

Why temperature and max tokens are core. They are the pair that most often goes missing. Across a large corpus of published evaluation results, a basic reproducibility field was absent from 96.5 percent of them, with max tokens missing from 95.6 percent and temperature from 93.9. Neither can be reconstructed from a result. The rest of the core is the smaller set without which a reader cannot say what was measured on what: the subject and its version, the claim, the data and its size, the metric and its direction, the number of runs, the headline, the limits, and who ran it and when.

Why a gap has to carry a reason. A blank field and a field nobody could fill look identical in a document, and the reader cannot tell which one they are looking at. So a null on this card is refused unless something beside it says why. "No seed was set" and "we did not write the seed down" are different facts about your run, and only one of them is a problem.

Why the limits field is core. Of 7,433 dataset cards studied on one public hub, the limitations section was the one most often left for last and most often left empty. It is also the section a second reader most needs, because it is the only place a card says where its number stops applying.

Why a seed still earns a field when the temperature is pinned. Setting temperature to zero is not the same as making a run reproducible. In one study of judged safety evaluations across 690 calls, two providers and three tiers, one or two of seven borderline items still came back differently under forced greedy decoding. So the settings that pin a run are recorded separately from the temperature: the seed, and, where a model did the judging, the version of the prompt it was judged with.

The digest. A data source named in words can change under a number without anyone noticing. A SHA-256 over the bytes that were actually scored cannot. That is why the field is here and why it wants the digest of the scored file rather than of a directory.

Formula: recorded = fields with a value · core recorded = core fields with a value

Questions

Questions

Is this the EvalEval schema?

No. It is aligned with it at the concept level and dated, which is what the note above the schema block records. We take the four signals as vocabulary and the minimal reproducibility pair in meaning, and we pin our own field list. Mirroring a schema still in beta field for field would mean every upstream rename silently invalidates cards people already exported.

What happens when v2 ships?

It ships beside v1, at its own URL, and v1 stays exactly where it is. A card you export today keeps validating forever. Inside v1 we may add optional fields; a required field, a rename, a removal or a narrowed type all publish a v2 instead. The version lives in the path, so nothing you already have has to be migrated.

Why is there no score?

Because a single number over completeness would be read as a quality rating of the evaluation, and it is not one. A well-documented weak evaluation would score higher than a thin note about a careful one. The counts tell you how many facts travel with the number; the verdict line tells you whether the core is there. Neither ranks anything.

Can I edit the JSON by hand?

Yes, and you can paste it back in. It validates against the published schema, so any validator will tell you whether your edit is still a v1 card. Two rules the schema cannot state are checked here instead: every null field needs a reason beside it, and a field you did record cannot also be listed as unrecorded.

What happens to a card written by a later version?

It depends on what "later" means, and the difference is worth knowing before you paste. A card that still says v1 and carries extra fields loads: the schema accepts unknown fields at every level, and this page names the ones it does not recognize instead of passing over them in silence. It does not carry them back out, so re-exporting from this form drops them. A card that declares a later version in schemaVersion is refused, and says which version it is and which one this page reads. That is deliberate rather than a gap: a v1 reader that guessed at v2 fields would hand you a card that looks read and is not.

Can I leave fields blank?

Yes, and the card still exports. A partial card that names its own gaps is the point of the page. What will not export is a card whose gaps are silent: leave a field empty and the box beside it empty too, and the export stops and names the field.

Sources

  1. Evaluation Cards: An Interpretive Layer for AI Evaluation ReportingarXiv preprint Retrieved
  2. Navigating Dataset Documentations in AI: A Large-Scale Analysis of Dataset Cards on Hugging FacearXiv preprint (ICLR 2024) Retrieved
  3. Croissant: a metadata format for machine-learning datasetsMLCommons Retrieved
  4. Necessary but Not Sufficient: Temperature Control and Reproducibility in LLM-as-Judge Safety EvaluationsarXiv preprint Retrieved