Blog
Writing on evaluation systems, judge architecture, workflow governance, and how teams ship AI with less guesswork.
How to Turn Production Agent Failures into Regression Tests
Capture one attributable failure, preserve its evidence, and turn it into a reviewed regression case without blindly promoting production data.
Why Aggregate AI Eval Scores Hide Regressions
Averages can improve while a critical behavior gets worse. Protected slices make that failure visible before release.
How to Test an MCP Server Before You Ship It
An MCP handshake is only the beginning. Test discovery, auth, schemas, isolation, failure semantics, and real client behavior.
How to Evaluate an AI Content Agent Before It Can Publish
Let the agent research and propose, but keep publishing behind tested code, explicit approval, and evidence you can measure later.
Why Every AI Product Needs Evaluation
Why agent workflows, judge evidence, trajectory quality, and enterprise controls have to live inside the same evaluation system.
Building Effective LLM Judge Systems
Rubrics matter, but strong judging depends on registry-backed models, disagreement handling, reliability, and evidence presentation.
Tracing: The Missing Layer in LLM Observability
Follow an AI request across retrieval, model, and tool steps to debug quality, latency, and cost with attributable evidence.
The Evolution of AI Testing: From Unit Tests to A/B Tests
A lifecycle view of evaluation methods, from deterministic checks to production experiments.
How to Turn Support Chatbot Failures into Release Evidence
An illustrative, evidence-first playbook for converting recurring support-agent failures into safer release gates.
Human-in-the-Loop: When to Use Annotations vs LLM Judges
A practical framework for deciding between human review, judge orchestration, and hybrid evaluation loops.