“What frameworks exist for evaluating AI agent behavior and safety: behavioral regression testing, adversarial evaluation, trajectory evaluation? Include METR, Apollo Research, Haize Labs, UK AI Safety Institute evaluations, and open-source eval harnesses.”
This landscape covers 22 representative frameworks, methods and organizations drawn from a curated survey of 25 projects plus focused candidate sets for METR, Apollo Research, Haize Labs, the UK AI Security Institute and 46 open-source projects. It is a representative survey spanning many independent hosts, not an exhaustive universe; coverage marks below reflect only what each source states, and blank cells mean the source did not claim support.
Each row is one framework or method; filled squares mark the evaluation modes its source documents. Hover or tap a square for the stated evidence.
METR's Task Standard gives tasks a common format run on agents via its Vivaria sandbox infrastructure, scored by verification scripts and calibrated against human baselines. Its time-horizon evaluations run models on 170 tasks from the HCAST, RE-Bench and SWAA suites — RE-Bench covering expert-level AI R&D work — each with human time-to-complete estimates, to estimate 50% and 80% time horizons. metr.org
Apollo runs agentic scheming evaluations (oversight subversion, self-exfiltration, goal-guarding, sandbagging) with ~300 rollouts per model, counting covert actions only after manual review of actions and chain-of-thought. Crucially, it tests evaluation awareness itself: editing evaluation-aware reasoning out of o3's traces raised its covert action rate from 13.2% to 24.2% — benchmark compliance is not alignment. alphaxiv.org
Cascade replaces human red-teamers with attacker LLMs running a tree/beam search over parallel conversation branches, scored by automated judging pipelines, to find multi-turn trajectories where benign queries escalate into jailbreaks. Verdict, its open-source judge library, stacks judge-then-verify LLM pipelines to make rubric scoring reliable. haizelabs.com
Inspect is AISI's open framework for model and agent evaluations with sandboxed environments hardened for adversarial agents; Inspect Evals adds community benchmarks including AgentDojo, AgentHarm and StrongREJECT; Transect analyzes long-horizon agent transcripts with scanners and turn-based timelines; safeguard evaluations feed red-team results into a misuse uplift model. github.com
Limitations: LLM judges need calibration and spot checks; hidden chain-of-thought access is often unavailable, limiting deception analysis; benchmarks contaminate and saturate; realistic long-horizon evaluations are expensive and highly scaffold-sensitive.
| Framework | Owner | Focus | Availability | Source |
|---|
Method: curated landscape of 25 agent-evaluation projects (one row per normalized project, multi-host survey), plus focused candidate sets for METR/Apollo (80 rows), Haize/UK AISI (100 rows) and open-source harnesses (46 deduplicated projects from 135 rows), collected 2025–2026. The matrix marks an evaluation mode only when the cited source states support; blank means unstated, not absent. Duplicate reports and paper-only narrow benchmarks were cut for space.