INSTRUMENT | Reliability testing
Context-of-Use Profiler: Turn a System Description into Evaluation Consequences
Describe what your AI system does, who it affects and how it fails, and get the failure families, case classes, slices and evidence expectations those answers demand.
Answer eleven questions about what the system does, who it affects and how its answers can be checked. The profiler returns the failure families, case classes, slices, evaluator constraints, evidence expectations and acceptance considerations those answers call for. Built on our agent-reliability research. Feed the profile into the blueprint builder. To start from specific failure concerns instead of a system description, use the risk-to-test mapper.
Showing your last valid result. Update the inputs above to recompute.
Scope
Unresolved
These dimensions are still marked "not known yet." Each entry names what answering the dimension would settle, so the unresolved list is a task list.
Likely failure families
Required case classes
Important slices
Evaluator considerations
Evidence expectations
Acceptance considerations
Dimensions that mattered
These answers changed at least one item in the profile above.
Dimensions that changed nothing
Recorded, but on their own they did not add or remove anything.
Rule trace
| Rule | Fired | Stopped at |
|---|
How the profile is built
No arithmetic. Every item in the profile is written by a named rule, and a rule fires when every one of its clauses matches the dimension answers on the form. A clause is a plain test against one or two dimensions; the rule's verdict is a written sentence, never a number. The trace table shows which rules fired and, for those that did not, which clause stopped them.
Why "not known yet" is a real answer. A dimension left at "not known yet" does not default to anything. It produces an unresolved entry that names what the answer would have settled, so the reader knows what decision is still open. This is the point of the profiler: it turns "we have not decided" into "here is what we have not decided, and here is what it costs."
The dimension split. After all rules have run, each dimension is placed in one of two lists: dimensions that mattered (at least one fired rule read this answer) and dimensions that changed nothing (recorded, but no rule needed it). A dimension at "not known yet" appears in neither list; it appears under Unresolved instead.
The six buckets. Every fired rule writes its item into exactly one of six buckets: likely failure families, required case classes, important slices, evaluator considerations, evidence expectations, and acceptance considerations. The buckets are the evaluation-design consequences the profiler exists to surface.
Questions
What do I do with this profile?
Export it and feed it into the evaluation blueprint builder. The profile fills the system name and pre-fills the slices list and the threats list from its consequences; the claims, evaluators and run plan you write yourself. Or use it as a checklist: walk the six buckets and mark off every item your current evaluation already covers.
Why is there no score?
Because a context-of-use profile is a list of things to check, not a rating of the system. A count of fired rules would tell you how many items landed in the profile; it would not tell you whether any of them matters more than any other, and it would reward answering more dimensions regardless of whether the answers were right.
Why does every dimension offer 'not known yet'?
Because a profile that defaults silently is worse than one that says where it is uncertain. Leaving a dimension unanswered does not fill it with a safe choice; it surfaces what the answer would have settled, so the reader can go and find out. A profile with five unresolved dimensions is a profile that knows five things it does not know.
Can I change my answers and rebuild?
Yes. Change any value and the profile rebuilds on every keystroke. The result region is marked stale while you are mid-edit, and it clears the moment the new profile is ready.