Tag: LLM-as-a-judge

One Prompt Change, Twelve New Failures
A cleaner AI prompt improved latency, cost, and overall pass rate—while quietly creating twelve new failures. This practical case study shows how prompt regressions hide behind aggregate metrics, how to build behaviour-based regression tests, and how before-and-after evaluation results can prevent a seemingly better agent from reaching production.

The AI Agent Test Manual Is Now Available
The AI Agent Test Manual is a practical guide to evaluating, testing, monitoring, and trusting production AI agents. Learn how to test prompts, tools, memory, retrieval, multi-agent workflows, LLM judges, release pipelines, observability, incident response, and the complete system surrounding an agent.

