Long-running AI agent workflows fail differently than simple background jobs: LLM calls are slow and expensive, outputs are non-deterministic, side effects like sending emails or charging cards can't be safely repeated, and some steps wait on humans for days. Durable execution solves this by checkpointing completed steps so failures resume from the last successful point rather than restarting entirely. Inngest is used as a case study, contrasted with BullMQ (queue-first, more operational control) and Temporal (workflow-first, more complex). Key building blocks covered include checkpointing via step.run(), waiting via step.sleep() and step.waitForEvent(), retry policies, idempotency for side effects, flow control (concurrency, throttling, rate limiting, debouncing, priority), and observability through tracing and replay. Developers still must handle idempotent side effects, step boundary design, retry policies, and cross-system data consistency themselves.

•31m read time•From newsletter.systemdesign.one
Post cover image
Table of contents
What Is an AI AgentWhy AI Agents Need Durable ExecutionWhy Traditional Job Queues Are NOT EnoughWhat Durable Execution Actually MeansHow to Build Durable Execution YourselfWhat Inngest IsBullMQ vs Temporal vs InngestUnder the Hood: How Inngest Executes a FunctionHow Inngest Handles FailuresOther Pieces That Matter in ProductionWhat You Still Have to HandlePutting It All Together: A Durable Research AgentClosing Thoughts

Questions this post answers

How is durable execution different from just retrying a failed job?

A retry reruns an entire failed task or workflow from the start, while durable execution (checkpointing) remembers which steps already completed successfully and only reruns the failed step. For a multi-step AI agent that retrieved documents, extracted evidence, and then failed during drafting, durable execution skips the retrieval and extraction and resumes at drafting, saving time, tokens, and avoiding duplicate side effects like re-sent emails. daily.dev surfaces this kind of workflow-recovery reasoning for engineers designing resilient agent pipelines.

When should I choose BullMQ, Temporal, or Inngest for building a durable AI agent workflow?

BullMQ fits teams with a dedicated infrastructure team who want fine-grained control over queues, workers, and retries but are willing to build checkpointing and resumability themselves. Temporal suits highly complex workflow graphs with many interdependent workflows, at the cost of added operational complexity. Inngest keeps workflow logic in application code via step.run(), handling checkpointing, retries, and waiting automatically, trading some low-level control for simplicity. developers weighing queue-first versus workflow-first tools for agent infrastructure can track these tradeoffs on daily.dev.

How does step.waitForEvent() let an AI agent workflow pause for human approval without wasting compute?

step.waitForEvent() in Inngest pauses a workflow's execution and releases its compute while waiting for a matching external event, such as an editor approving a draft, instead of holding a worker process open. A timeout can be configured for cases where the event never arrives, and once the event lands, the workflow resumes with its prior progress intact, even after waiting for days. teams building human-in-the-loop agent approvals can compare these waiting patterns on daily.dev.

1 Comment
Share this post