When an agent system fails in production — and given enough time, every agent system fails in production — it almost always fails in one of three ways. It loops forever, burning tokens and accomplishing nothing. Or its plan was wrong from the start and every well-executed step takes it further from the goal. Or its tool calls did real damage, because the tool surface was wider than the trust model warranted. The good news: each failure mode has a known root cause, and each root cause has a known engineering fix. This field guide is the diagnostic manual.

Companion to the Scaling Agentic AI and Production Playbook series. Those documents argue for verification as the missing architectural piece; this one is the failure taxonomy that motivates the argument.

What it covers

Three failure modes plus a checklist, about eleven minutes.

§ 1 — Failure 1: the infinite loop. Symptom: token cost climbs without progress. Root cause: the agent has no termination condition that the verifier actually enforces. Fix: explicit loop budgets, progress-checking verifiers, the did the state change check before the next iteration.

§ 2 — Failure 2: the broken plan. Symptom: every step is correctly executed and the result is wrong. Root cause: the plan was built before the agent had enough context; subsequent execution doesn’t notice the plan is failing. Fix: re-planning checkpoints, evidence-driven plan revision, the is this plan still working check between sub-tasks.

§ 3 — Failure 3: unsafe tool use. Symptom: the agent does real harm — leaks secrets, mass-deletes data, sends bad emails. Root cause: the tool surface was wider than the principle of least privilege. Fix: scoped tool calls, mandatory human-in-the-loop confirmation for irreversible actions, the minimum capability needed for the task discipline.

§ 4 — The diagnostic checklist. Twelve symptoms grouped by failure mode. When you see this in your logs, look here in your code.

Read it

Open the field guide →

The field guide lives at its own URL with a warm-paper layout — fault-red accent for symptoms, verify-green accent for fixes. Diagnostic manual format throughout.


← Back to Autonomy