Durable agents run on a layered runtime — a log, checkpoints, disposable workers, retry machinery — not on any single framework

Asked (summary)

What infrastructure do engineering teams build internally for durable AI agent execution — checkpoints, resumability, retries, idempotency, crash recovery? What open-source options exist and where do they fall short? Name the sources.

15 source-backed systems: 6 internal production builds described firsthand by Anthropic, Cursor, Cloudflare, and OpenAI engineering teams, and 9 open-source options drawn from vendor documentation and independent analyses, each scored against the seven runtime layers that recur across the production builds.

Who covers which layer

Layer evidenced in an internal production build (firsthand engineering account) Layer evidenced in an open-source option (docs or independent analysis) Not evidenced in the cited source Hover or tap a cell for the evidence
Anthropic's Managed Agents keeps the append-only session log outside the harness and sandbox, so a crashed harness is simply rebooted and resumes from the last event — anthropic.com
Cursor abandoned its home-grown work-stealing runtime ("one 9 of reliability") for Temporal plus durable VM checkpoint/restore and retry-aware stream rewind — cursor.com
Replay engines memoize completed steps, but that does not make external side effects exactly-once: Dapr's own analysis says durable execution is necessary but not sufficient without application-level idempotency — bit.ly
Self-hosting costs are real: Flyte requires Kubernetes, Hatchet pulls in RabbitMQ at throughput, DBOS ties state to Postgres, and Inngest caps run state at 32 MiB with 4 MiB per step — inngest.com

All 15 systems, with recovery behavior and limits

SystemCategoryDurability & recoveryWhere it falls shortSource

Capability matrix and table built from 15 source-backed rows (6 internal production builds, 9 open-source options), one per normalized team or project, scanned from live web sources on 2026-10-08. A filled cell means the cited source explicitly describes that runtime layer; an empty cell means the source does not mention it, not that the system lacks it. Architecture, recovery, and limitation texts are truncated to ~320 characters for space; full detail is at each linked source. Firsthand engineering accounts (blue) are distinguished from vendor documentation and independent analyses (teal).

This report was generated automatically by Keenable SELECT at a user's request, from publicly available web sources linked herein. Keenable does not review, verify, or endorse its contents and makes no representation as to accuracy, completeness, or timeliness; AI-based extraction may contain errors. Nothing in this report is investment, legal, financial, or other professional advice. All trademarks and referenced content remain the property of their respective owners; no affiliation or endorsement is implied. To report an error, rights concern, or request removal: legal@keenable.ai.

Keenable SELECTAsk your own question
Made with Keenable SELECT