Your Agent Passed the Demo. It's Still Failing in Production. You Just Can't See It Yet.
Something changed in how AI gets built — and almost nobody is measuring it correctly.
Your agent works in the demo. The refund processes, the ticket closes, the room nods. Then it ships. And somewhere out in production it's quietly telling customers the wrong policy, burning fifty dollars on a task that should cost forty cents, and leaking one customer's data to another — all while every dashboard glows green. No crash. No error. No alert. Just silent, expensive, trust-destroying failure that you won't discover until a customer does.
Here's the uncomfortable truth: evaluating an AI agent is nothing like evaluating a chatbot. A chatbot returns a sentence you can eyeball. An agent produces a trajectory — a chain of tool calls, decisions, and actions where any single step can go wrong in ways a spot-check will never catch. The old playbook of "run it a few times and see if it looks good" isn't just weak. It's the reason your agent is failing right now without your knowledge.
This book takes you from "I hope it works" to "I can prove it works — and prove it stays working."
You'll build a complete evaluation program from the ground up, following one realistic system the entire way: Northwind Support, a customer-service agent with real tools, real policies, and real failure modes you'll learn to catch by construction. This isn't theory. It's a working project you build chapter by chapter — instrumented traces, a coverage-driven test dataset, code-first graders, calibrated LLM judges, and a CI gate that blocks bad changes before they reach a customer.
By the end, you will know how to:
The window to build this discipline is open right now. A small group of engineers is already treating evals as the core skill of the AI era — gating every deploy, catching regressions in minutes instead of incidents, and shipping agents they can actually defend to a VP, an auditor, or a customer. Everyone else is still running their agent on faith and hoping.
Written in a practitioner-to-practitioner voice — no filler, no academic detachment, no AI-sounding fluff — every technique is shown in runnable Python against the same system, so you learn operationally instead of abstractly. Complete appendices give you a 2026 tooling field guide, a critical benchmark reference, copy-ready judge prompts and rubric templates, and a full environment setup so every listing runs as written.
Your agent is already failing quietly. The only question is whether you catch it first — or your customers do.
Stop shipping on hope. Start shipping on measurement. Build the system that makes your agents provably reliable — starting today.
Die Inhaltsangabe kann sich auf eine andere Ausgabe dieses Titels beziehen.
Anbieter: California Books, Miami, FL, USA
Zustand: New. Print on Demand. Bestandsnummer des Verkäufers I-9798191291697
Anzahl: Mehr als 20 verfügbar
Anbieter: PBShop.store UK, Fairford, GLOS, Vereinigtes Königreich
PAP. Zustand: New. New Book. Shipped from UK. Established seller since 2000. Bestandsnummer des Verkäufers L2-9798191291697
Anzahl: Mehr als 20 verfügbar
Anbieter: AHA-BUCH GmbH, Einbeck, Deutschland
Taschenbuch. Zustand: Neu. Neuware. Bestandsnummer des Verkäufers 9798191291697
Anzahl: 2 verfügbar