Most Gen AI applications do not fail with an error message. They fail while sounding convincing. This practitioner's playbook shows engineering, product, testing, risk and governance teams how to build evaluation systems that expose those failures before they reach users, auditors or regulators. Using the running Meridian case, the book takes you from the first evaluation test to a governed production operating model.
It shows how to create versioned datasets, evaluation rubrics, traces, regression suites, human-review workflows, release gates, model-change decisions, red-team evidence and audit-ready release records. Inside the Second Edition Build an evaluation suite from business requirements, failure modes and protected behaviours. Evaluate RAG pipelines, agentic workflows, multimodal applications, prompts, code and reasoning.
Govern model migrations and fine-tuning changes using non-compensatory release gates. Apply uncertainty analysis, retrieval metrics, pass@k, human calibration and defensible thresholds. Test MCP and tool-use boundaries, memory failures, prompt injection, provenance and adversarial behaviour. Connect evaluation evidence to OWASP, MITRE ATLAS, C2PA, the EU AI Act and production governance. Practise the methods through seven interactive companion applications built around realistic Meridian scenarios.
This is not a prompt book or a catalogue of benchmark scores. It is an engineering guide to deciding what to test, what evidence to retain, when to stop a release and how to keep evaluation effective after deployment. Written for AI engineers, software testers, architects, platform teams, product leaders, risk professionals and AI-governance practitioners responsible for Gen AI systems in production.
Most Gen AI applications do not fail with an error message. They fail while sounding convincing. This practitioner's playbook shows engineering, product, testing, risk and governance teams how to build evaluation systems that expose those failures before they reach users, auditors or regulators. Using the running Meridian case, the book takes you from the first evaluation test to a governed production operating model.
It shows how to create versioned datasets, evaluation rubrics, traces, regression suites, human-review workflows, release gates, model-change decisions, red-team evidence and audit-ready release records. Inside the Second Edition Build an evaluation suite from business requirements, failure modes and protected behaviours. Evaluate RAG pipelines, agentic workflows, multimodal applications, prompts, code and reasoning.
Govern model migrations and fine-tuning changes using non-compensatory release gates. Apply uncertainty analysis, retrieval metrics, pass@k, human calibration and defensible thresholds. Test MCP and tool-use boundaries, memory failures, prompt injection, provenance and adversarial behaviour. Connect evaluation evidence to OWASP, MITRE ATLAS, C2PA, the EU AI Act and production governance. Practise the methods through seven interactive companion applications built around realistic Meridian scenarios.
This is not a prompt book or a catalogue of benchmark scores. It is an engineering guide to deciding what to test, what evidence to retain, when to stop a release and how to keep evaluation effective after deployment. Written for AI engineers, software testers, architects, platform teams, product leaders, risk professionals and AI-governance practitioners responsible for Gen AI systems in production.