LatentEval

INSTRUMENT | Reliability testing

Context-of-Use Profiler: Turn a System Description into Evaluation Consequences

Describe what your AI system does, who it affects and how it fails, and get the failure families, case classes, slices and evidence expectations those answers demand.

Answer eleven questions about what the system does, who it affects and how its answers can be checked. The profiler returns the failure families, case classes, slices, evaluator constraints, evidence expectations and acceptance considerations those answers call for. Built on our agent-reliability research. Feed the profile into the blueprint builder. To start from specific failure concerns instead of a system description, use the risk-to-test mapper.

The system

Used in the exported title. Optional, but a named profile is easier to find later.

Check this value.

Scope

What the system is for, in the operator's own words. This is what the evaluation has to cover.

Check this value.

Uses a reasonable person would try that the system is not meant for. These become case classes.

Check this value.

Uses the system should refuse outright. The boundary between foreseeable misuse and out-of-scope is what the refusal-boundary cases test.

Check this value.

Eleven dimensions

Check this value.

Check this value.

Check this value.

Check this value.

Check this value.

Check this value.

Check this value.

Check this value.

Check this value.

Check this value.

Check this value.

Notes

Context that does not fit a dimension. Carried in the export, not used by the rules.

Check this value.

Scope

Unresolved

These dimensions are still marked "not known yet." Each entry names what answering the dimension would settle, so the unresolved list is a task list.

Dimensions that mattered

These answers changed at least one item in the profile above.

    Dimensions that changed nothing

    Recorded, but on their own they did not add or remove anything.

      Rule trace
      Rule Fired Stopped at
      Export the profile
      How the profile is built

      No arithmetic. Every item in the profile is written by a named rule, and a rule fires when every one of its clauses matches the dimension answers on the form. A clause is a plain test against one or two dimensions; the rule's verdict is a written sentence, never a number. The trace table shows which rules fired and, for those that did not, which clause stopped them.

      Why "not known yet" is a real answer. A dimension left at "not known yet" does not default to anything. It produces an unresolved entry that names what the answer would have settled, so the reader knows what decision is still open. This is the point of the profiler: it turns "we have not decided" into "here is what we have not decided, and here is what it costs."

      The dimension split. After all rules have run, each dimension is placed in one of two lists: dimensions that mattered (at least one fired rule read this answer) and dimensions that changed nothing (recorded, but no rule needed it). A dimension at "not known yet" appears in neither list; it appears under Unresolved instead.

      The six buckets. Every fired rule writes its item into exactly one of six buckets: likely failure families, required case classes, important slices, evaluator considerations, evidence expectations, and acceptance considerations. The buckets are the evaluation-design consequences the profiler exists to surface.

      Questions

      What do I do with this profile?

      Export it and feed it into the evaluation blueprint builder. The profile fills the system name and pre-fills the slices list and the threats list from its consequences; the claims, evaluators and run plan you write yourself. Or use it as a checklist: walk the six buckets and mark off every item your current evaluation already covers.

      Why is there no score?

      Because a context-of-use profile is a list of things to check, not a rating of the system. A count of fired rules would tell you how many items landed in the profile; it would not tell you whether any of them matters more than any other, and it would reward answering more dimensions regardless of whether the answers were right.

      Why does every dimension offer 'not known yet'?

      Because a profile that defaults silently is worse than one that says where it is uncertain. Leaving a dimension unanswered does not fill it with a safe choice; it surfaces what the answer would have settled, so the reader can go and find out. A profile with five unresolved dimensions is a profile that knows five things it does not know.

      Can I change my answers and rebuild?

      Yes. Change any value and the profile rebuilds on every keystroke. The result region is marked stale while you are mid-edit, and it clears the moment the new profile is ready.