What are the unresolved system-level problems in AI agent memory and context management for production agents — persistent state, memory concurrency, context-window costs, context freshness, and memory provenance? Name the sources.
This synthesis draws on 95 sourced findings from a broad live-web scan of research papers, surveys and engineering discussions current through October 2026 — mostly primary arXiv papers, with mirrors and paper-list aggregators de-emphasized. It is a structured synthesis of unresolved issues, not an exhaustive census of the literature. Empirical results are labelled as such; surveys and proposed architectures are labelled separately, and no single benchmark number should be read as universal prevalence.
Aggressive compression saves tokens but severs lineage and freshness: a summary that drops a revocation or a rare qualifier is cheaper and wrong. Concurrency multiplies lineage: every derived copy an agent holds must be reachable by a later deletion or revocation. Persistence turns one bad write into a cross-session failure that retrieval keeps resurfacing. And every remedy — verification, provenance capture, event sourcing — adds context, latency and storage cost back onto the window the agent was trying to shrink. The measured safety–utility frontier in authorization laundering (unauthorized use cut from 25.3% to 7.3%, authorized use falling to 53.8%) is the clearest evidence that these are coupled trade-offs, not independent bugs.
| Area | Unresolved system problem | Why current approaches fall short | Production consequence | Partial direction | Source |
|---|
Structured synthesis of 95 sourced findings (30 on persistent state, 25 on concurrency and shared state, 40 on context cost, freshness and provenance) gathered by broad live-web search across papers, surveys, framework docs and engineering discussions, current through 2026-10-08. Duplicate versions of the same paper were merged and aggregator rows de-emphasized in favour of primary papers. Percentages are each paper's own reported benchmark results, not population prevalence. The pipeline map shows 23 representative failure modes; the table condenses to 13 problem rows for space. This is not an exhaustive census of the literature.