Production AI agents fail differently than single requests because runs are long, stateful, and take real-world actions. The piece lays out five resilience patterns: retry the failed step rather than the whole run and sort failures into time-fixable, model-fixable, or unfixable categories; make every write tool idempotent with a stable key; checkpoint state after each step so crashes mean resume rather than restart; fall back to the same model on a different cloud provider rather than a different model to avoid shared outages and eval drift; and bound every run by steps, tokens, and wall-clock time, escalating stuck runs to a human. Code examples use Python with the Inngest durable-execution SDK to implement step-level retries, checkpointing, and human-in-the-loop approval waits.

•8m read time•From thetshaped.dev
Post cover image
Table of contents
An agent isn't a request1. Retry the step, not the run2. Make every tool safe to call twice3. Checkpoint, so a crash means resume4. Fall back across failure domains, not just models5. Bound the loop, then hand it to a human📌 TL;DR

Questions this post answers

Why do AI agent runs fail so often even when each model or tool call succeeds 98% of the time?

Failure compounds across steps: at 98% success per call, only about 67% of a 20-step agent run completes without a single failure, meaning roughly one in three runs hits a failure. Even at 99% success per call, one in five 20-step runs fails, which is why failure has to be treated as the main path rather than an edge case in agent design. Anyone architecting multi-step agent workflows can track failure-handling patterns like these through daily.dev.

How do I make a tool call safe to retry in an AI agent workflow?

Give every tool that writes data an idempotency key built from stable IDs, such as a model's tool-call ID, rather than a key generated fresh inside the step. Stripe popularized this pattern for payment APIs and keeps idempotency keys for 24 hours; applying the same approach lets a retried step re-run safely without duplicating side effects like refunds or ticket creation. Developers wiring up idempotent tool calls can keep up with patterns like this via daily.dev.

What is the difference between retrying an entire AI agent run versus retrying just a failed step?

Retrying the whole run re-executes and re-pays for every step that already succeeded, replaying their side effects like duplicate emails or refunds, while retrying only the failed step preserves completed work. Failures should be sorted into three categories: time-fixable errors like 429s or timeouts get retried later, model-fixable errors get sent back as a tool result for the model to correct, and unfixable errors like a revoked API key should fail the step immediately. Teams debugging flaky agent retries can follow step-level failure-handling approaches on daily.dev.

Share this post