“What open-source tools exist in 2026 for monitoring autonomous AI agents at runtime: observing tool calls, detecting anomalous behavior, enforcing policies, intervening before side effects? Name specific projects and compare their approaches.”
Research against live web data on 8 October 2026 surfaced 32 named open-source projects — 24 runtime-control or containment projects and 8 primarily observability projects. They split into three complementary layers: observability records traces but generally cannot stop an action; in-process guardrails and MCP/AI gateways inspect tool calls and can allow, deny, redact, rate-limit, pause for approval, or kill a run; OS-level sandboxes contain code after a call is admitted. This is a practical landscape, not an exhaustive or maturity-ranked universe.
All 32 projects, grouped by where they sit and what they can do before a side effect. Tile width reflects research coverage (pages found mentioning the project in this research — not popularity or quality). Hover or tap a tile for its enforcement point and controls.
Capabilities read from the collected source text only. A filled square means the source describes the capability; an empty square means not stated in the collected source — it does not mean the project lacks it. Observability projects are scored as telemetry layers.
The layers complement rather than replace each other; a reliable deployment typically combines all three.
OpenLLMetry, MLflow, OpenLIT, SigNoz, Arize Phoenix, Langfuse, AgentOps, and Helicone record prompts, tool calls, and costs over OTLP or compatible SDKs. Tracing is retrospective: it explains an incident but does not prevent one.
Pre-side-effect intervention needs an enforcement point the agent cannot route around — an MCP/AI gateway (agentgateway, Portkey, ToolHive, Kong, fast-gateway) or an in-process interceptor (NeMo Guardrails, LlamaFirewall, AgentTrust, APort). Detection ranges from deterministic rules, quotas, and allowlists (fast, low false positives) to classifiers, LLM judges, and multi-step behavioral analysis (broader coverage, higher latency and false-positive cost).
E2B, Daytona, Sandlock, DSec, and the Sandbox Platform restrict files, network, processes, and credentials once a call is admitted — Firecracker microVMs, gVisor pods, Landlock/seccomp, eBPF egress filters. Pair with least-privilege credentials so a compromised agent gains little.
Search or filter by layer; click a project name for the full source text on approach, detections, controls, and enforcement point. Blank fields mean the collected source did not state it.
| Project | Layer | Pre-side-effect control (summary) | Pages | License | Source |
|---|
Method: comparison of 32 open-source projects for runtime monitoring and control of autonomous AI agents, one row per project, compiled from live web research on 2026-10-08. “Pages” counts distinct relevant pages found in the candidate research (coverage, not popularity or quality). Capability squares are derived only from the collected source text; an empty square means not stated, not unsupported. Licenses are blank where the source did not establish one; several newer projects are documented by papers and should be validated for repository activity and maturity before deployment. Long source passages are truncated in the expanded rows only where the source itself was cut.