Evaluating AI agents in production
What breaks when agents meet real tools, and how to test for it before launch.
Insights
Three series are in preparation. Articles appear here as they are published.
What breaks when agents meet real tools, and how to test for it before launch.
How to read an eval report, and the questions that expose a weak one.
Why most vendor-built frameworks die within a year, and how to build one that does not.
The campaign structure, harm taxonomy and retest loop we run before any launch.
Retrieval precision, answer faithfulness, and the gates that keep a pipeline honest.
Building load models from production traffic, and fixing what they expose in time.