Guides
Written for two readers: the engineer shipping agents, and anyone deciding how far to trust the AI they use.
For builders
Model choice, testing, observability, and security for engineers shipping agents.
Evals and accuracy
-
Evaluation case schema, field by field
The case profile published atop the eval-set schema: what each field does, why the oracle type gates the evaluator, what is left out on purpose, and how rows round-trip.
-
Evidence quality for AI claims: eleven dimensions, no score
Eleven dimensions for appraising evaluation evidence, named rate-down reasons for each, and a five-part argument against collapsing them into a composite number.
-
Golden set blueprint: coding agents
Seven case families for coding agents: must-flip and must-not-break tests on separate denominators, solution leakage measured, and the test suite treated as the evaluator it is.
-
Golden set blueprint: customer support
Seven case families, five slices, a coverage skeleton, and ten starter rows for testing a support agent, with a finance overlay and CSV/JSONL downloads.
-
Golden set blueprint: RAG and knowledge assistants
Seven case families for a RAG assistant, five slice families, oracle types and evaluator guidance per family, downloadable starter rows, and a pharma overlay.
-
Golden set blueprint: regulated content
Seven case families for content under regulatory constraints: per-output constraint checks, risk-disclosure preservation, audience mismatch, and a pharma overlay. Not compliance advice.
-
Golden set blueprint: research agents
Seven case families for research agents: citation existence, citation support, evidence strength, source coverage, and silent fabrication, each graded per claim against a named source.
-
Golden set blueprint: structured extraction
Six case families for document extraction, counting omission and hallucination separately and reporting both per-field and per-record metrics.
-
Golden set blueprint: summarization
Six case families for summarization: source contradictions, unsupported additions, and salient omissions counted separately, with evaluator granularity tested as a variable.
-
Golden set blueprint: transactional agents
Seven case families for agents that change world state: expected writes, confirmation integrity, silent tool failure, budget exhaustion, and a finance overlay.
-
OpenAI Evals is winding down. The alternatives skip the statistics.
OpenAI is deprecating its hosted Evals platform and steering users to Promptfoo, which it now owns. What OpenAI Evals, DeepEval, Ragas, TruLens, and Promptfoo each do, and what porting costs you.
-
AI observability proves the run finished. An eval proves it was right.
LangChain's State of Agent Engineering survey of 1,340 practitioners: 89% run observability, only 52.4% run offline evals. An offline eval is what scores whether the answer was right.
-
The prompt wording is a hyperparameter you never swept.
Rewording the same task swings a model's pass rate: format, option order, even a 'please'. A one-phrasing eval samples one point from a spread you never measured. Pin the prompt and measure it.
-
How benchmarks get gamed, and how to check yours
BenchJack, a Berkeley auditing tool, found 219 flaws in ten popular agent benchmarks and gamed nine to near-perfect scores. Why a benchmark number is a claim about the harness, and how to check it.
Model choice and routing
-
We scored Fable 5 vs Opus 5 vs Opus 4.8 and only the old one answered every call
While Claude Fable 5 and Claude Opus 5 are tied on our reliability benchmarks, Opus 4.8, surprisingly, is still the better choice for two specific kinds of work.
-
We scored Fable 5 vs Sol vs Kimi K3 and the winner refused the most calls
Claude Fable 5 wins on points, GPT-5.6 Sol comes last and answers everything, and Kimi K3 is quick until it hangs. Pick by the failure your pipeline can absorb.
-
The best Claude model for coding is rarely the one at the top.
Which Claude model should you use? A task-by-task routing guide across Opus 4.8, Sonnet 5, Haiku 4.5, and Fable 5, with a checkability rule for coding, writing, research, and agents.
-
Opus vs Sonnet: the routing question your task already answered
Claude Opus 4.8 lists at 1.67x Sonnet 5, yet on checkable work our routing eval measured no separation between them. When the two workhorse tiers are interchangeable, and when the premium earns out.
-
Buy down, not up: what reaching for the top Claude tier really costs
Is Fable 5's 2x worth it on work a script can grade? Our routing eval measured no Opus-to-Fable separation at either effort, and the premium bought a refusal tax and a quiet fallback to Opus.
Monitoring and tooling
-
Hamming vs Cekura: voice agent eval platforms compared
A source-checked Hamming vs Cekura comparison: published pricing against sales-call pricing, a self-refereed 90% win rate, judge agreement quoted without a base rate, and how to run your own bake-off.
-
Datadog LLM Observability vs Langfuse for agent evals
A source-verified comparison of Datadog Agent Observability (formerly LLM Observability) and Langfuse for agent evals: pricing, judge templates, self-hosting, and the statistics neither documents.
-
Helicone alternatives after the Mintlify acquisition
Helicone is in maintenance mode after the March 2026 Mintlify acquisition. Where its users go next, split by surface: gateway replacements for proxied traffic, observability platforms for logging.
-
Langfuse vs LangSmith: the real split, and what all three still miss
A verified comparison of Langfuse, LangSmith, and Braintrust on licensing, self-hosting, tracing, evals, and pricing shape, plus the statistical eval layer none of the three computes for you.
-
Model fallback swaps the model unseen, and refusals log as success
When Claude Fable 5 refuses, it returns HTTP 200 with an empty content array, so error dashboards read it as a success, and with fallback on the model can swap to Opus 4.8 unseen. Instrument now.
MCP and protocols
-
Every tool behind one MCP gateway. One breach reaches all of them.
An MCP gateway fronts many MCP servers as one endpoint, brokering credentials and policy. It buys one audit trail and concentrates every tool into one process. Gateway vs proxy, when to run one.
-
SEP-2567 deletes Mcp-Session-Id, and sticky routing with it
The 2026-07-28 spec shipped. SEP-2567 deleted Mcp-Session-Id and sticky routing with no grace period. Roots, Sampling, and Logging keep a year, but not every method. Audit the session layer first.
Security and privacy
For everyday use
Plain-language guides for people using AI tools without building them.
Consumer agents
-
Amazon v. Perplexity: the agent-blocking ruling that didn't survive appeal
You told your shopping agent yes, and Amazon wanted a court to say no; it won that order against Perplexity's Comet, and on August 4, 2026 the Ninth Circuit vacated it.
-
Is Chrome's agentic browser safe when it clicks Buy for you?
Google's auto browse lets an AI click and type across your tabs, in a US-only paid preview. The same agent obeys instructions hidden on the pages it reads. How to keep a hand on it.
-
Whose side is your AI shopping agent on?
Your AI shopping agent has your card and picks what you buy. Google, Amazon, and Perplexity all ship one now, and when a model carries a sponsor's incentive, it steers. Learn whose side yours is on.