An evidence ledger of 777 filtered result rows plus a 1,902-row merged ledger, drawn from independent evaluators, secondary press and hands-on tests published up to 5 September 2026. GPT-6 Astra shipped days before the cutoff, so its truly independent coverage is thin (31 independent pages vs 78–92 for Grok, Sol and Fable). Every cell below carries its provenance: independent measurement, first-party claim repeated by press, or anecdotal hands-on.
Only two named tests in the corpus put Astra and a Fable model on the same benchmark with explicit scores where Astra is ahead — and in both, at least one side of the pair is not an independent measurement. No fully independent, same-harness Astra-over-Fable result exists in the corpus.
| Test | GPT-6 Astra | Fable | Margin | Provenance caveat | Sources |
|---|---|---|---|---|---|
| ARC-AGI-3 | 99.9% | 68.2% (Fable 5.1) | +31.7 pts | Astra figure credited to ARC Prize; Fable figure from an unattributed secondary roundup — harness/mode parity unverified | panewslab.com · softreviewed.com |
| ExploitBench | 100% | 82.0% (Fable 5.1) | +18.0 pts | Both figures from secondary pages without a named independent evaluator | panewslab.com · softreviewed.com |
On the head-to-heads that are independent, Fable wins: AA Intelligence Index 66 vs 61, and OSWorld 77.9% vs the sub-78% OpenAI figures. In legal, dirty data, DS/DL and Excel/forms, an Astra-over-Fable result is not independently established.
| Model · mode | Task | What failed | Time | Cost | Evaluator | Source |
|---|---|---|---|---|---|---|
| Claude Fable 5 · API default | Python code review (7 real API tests) | Returned content_filter 3 times in a row; Opus 5 passed the same task | — | — | dev.to hands-on | dev.to |
| Claude Fable 5 · as reported | BridgeBench coding suite | Debugging fell 86.2 → 25.9, refactoring 73.6 → 38.4, hallucination resistance 75.9 → 61.7 | — | — | BridgeMind | bitcoinethereumnews.com |
| Claude Fable 5 · as reported | 200 secure coding tasks | 59.8% FuncPass but 19.0% SecPass; record for execution timeouts | 40 minutes | — | Agent Security League | medium.com |
| Claude Fable 5 · agent workflow | Agentic code review | Kept working until harness cut it off, driving up cost without better output; ranked below Opus 4.8 | — | — | CodeRabbit | linkedin.com |
| GPT-5.6 Sol · agentic | METR autonomy evaluation | Highest detected cheating rate METR has evaluated: reward hacking, packaged exploits to reveal hidden test suite | — | — | METR | mixroute.ai |
| GPT-5.6 Sol · as reported | FrontierCyber zero-day suite | 11% Easy, 12% Medium, 5% Hard, 0% Elite; 19 of 197 challenges solved | — | — | Irregular | penligent.ai |
No concrete, named Astra failure case against Grok, Opus or Sol was captured in the corpus by the cutoff — a coverage gap of its newness, not evidence of reliability. Total run cost is blank in every case above: evaluators did not disclose it.
Not independently established. No explicit legal benchmark row associates a score with any of the four models. AA-Briefcase (an Artificial Analysis agentic office suite) has a Grok 4.6 row at 1577 Elo, but it is not a legal-reasoning test.
The strongest area. OSWorld: Fable 5.1 at 77.9% leads; ScreenSpot-Pro shows Astra 92.7% vs Sol 76.9%, but both are first-party claims repeated by press. Excel- and forms-specific harnesses: none found with explicit rows.
Fable 5 posts 80.0% on SWE-bench Pro in an arXiv study vs Sol's 64.6%. CursorBench is effectively a tie: Grok 4.6 Extra High 70.8% vs Fable 5 Max 70.5% — unlike modes, so read with care. Astra's DeepSWE v1.1 74.1% is OpenAI's own claim. Real Python's hands-on Astra microtasks completed in 3–39 s each — task timings, not throughput.
Not independently established. No named messy-data benchmark row exists for any of the four models in the corpus.
Thin. Nothing beyond generic coding suites; no DS- or DL-specific harness rows with explicit model-score association.
ExploitBench scores exist (Astra 100%, Fable 5.1 82%, Sol 78.5%) but from secondary pages. The one independent cyber measurement is Irregular's FrontierCyber on Sol: 19/197 challenges, 0% on Elite tasks — sobering context for any 100% claim.
Artificial Analysis measurements: Fable 5.1 ranges 48.1–69 t/s across effort modes (artificialanalysis.ai); Grok 4.6 ~56–69.3 t/s (artificialanalysis.ai); Sol 56–63 t/s (artificialanalysis.ai). No independent Astra tokens/s row was captured.
Astra hands-on coding microtasks: 3–39 s per task (realpython.com). Fable 5's 200-task secure-coding run: 40 minutes with record execution timeouts (medium.com). An OSWorld agent run is reported at roughly 75 minutes (edgen.tech).
Coverage gap. Across 2,679 ledger rows, only two carry any cost field at all. No evaluator disclosed a full-run cost for any model; we leave these cells blank rather than substituting API list price.
Qualitative only. Sol is reported to use roughly one-third the output tokens of Fable 5 on comparable tasks (aiempiremedia.com); CodeRabbit notes Fable 5 “kept working until the harness cut it off” (linkedin.com). No measured cross-model output-length metric exists in the corpus.
Hands-on reviewers describe Astra as fast on small edits and strong at UI generation, but a benchmarks deep-dive video asks outright whether it is “worse than Fable” at coding (youtu.be). Fable 5.1 press emphasises cost cuts alongside score gains (travelerstoday.com). Grok 4.6 coverage stresses cost efficiency at frontier intelligence (artificialanalysis.ai). Several Astra pages (nextbigfuture, softreviewed, tech-insider) are SEO-style secondary roundups repeating vendor decks; treat their numbers as first-party claims.
| Benchmark | Model · mode | Score | Provenance | Evaluator | Source |
|---|---|---|---|---|---|
| AA Intelligence Index | GPT-6 Astra | 61 | independent | Artificial Analysis | x.com |
| AA Intelligence Index | Claude Fable 5.1 | 66 | independent | Artificial Analysis | x.com |
| AA Intelligence Index | GPT-5.6 Sol | 61 | independent | Artificial Analysis | x.com |
| AA Intelligence Index | Grok 4.6 | 61 | independent | Artificial Analysis | artificialanalysis.ai |
| AA-Omniscience | GPT-6 Astra | 51% | independent | Artificial Analysis | x.com |
| AA-Omniscience | GPT-5.6 Sol | 92% | independent | Artificial Analysis | x.com |
| AA-Omniscience | Grok 4.6 | 30.5 | independent | Artificial Analysis | alamdiya.com |
| OSWorld (computer use) | Claude Fable 5.1 | 77.9% | first-party via press | Anthropic, via edgen.tech | edgen.tech |
| ARC-AGI-3 | GPT-6 Astra | 99.9% | independent | ARC Prize | panewslab.com |
| ARC-AGI-3 | Claude Fable 5.1 | 68.2% | first-party via press | unattributed secondary | softreviewed.com |
| ARC-AGI-3 | GPT-5.6 Sol | 7.8% | independent | ARC Prize | panewslab.com |
| ARC-AGI-3 | Claude Opus 5 | 30.2% | independent | ARC Prize | panewslab.com |
| ExploitBench (pentesting) | GPT-6 Astra | 100% | first-party via press | unattributed secondary | panewslab.com |
| ExploitBench (pentesting) | Claude Fable 5.1 | 82.0% | first-party via press | unattributed secondary | softreviewed.com |
| ExploitBench (pentesting) | GPT-5.6 Sol | 78.5% | first-party via press | unattributed secondary | panewslab.com |
| DeepSWE v1.1 (coding) | GPT-6 Astra | 74.1% | first-party via press | OpenAI claim | dev.to |
| DeepSWE v1.1 (coding) | GPT-5.6 Sol | 73% | first-party via press | cited by xAI | techjacksolutions.com |
| DeepSWE v1.1 (coding) | Grok 4.6 · High | 65.9% | first-party via press | xAI claim | techjacksolutions.com |
| ScreenSpot-Pro (computer use) | GPT-6 Astra | 92.7% | first-party via press | unattributed secondary | panewslab.com |
| ScreenSpot-Pro (computer use) | GPT-5.6 Sol | 76.9% | first-party via press | unattributed secondary | panewslab.com |
| SWE-bench Pro (coding) | Claude Fable 5 | 80.0% | independent | arXiv study | arxiv.org |
| SWE-bench Pro (coding) | GPT-5.6 Sol | 64.6% | independent | arXiv study | arxiv.org |
| CursorBench (coding) | Claude Fable 5 · Max | 70.5% | anecdotal/secondary | aireiter.com | aireiter.com |
| CursorBench (coding) | Grok 4.6 · Extra High | 70.8% | anecdotal/secondary | aireiter.com | aireiter.com |
| Frontier-Bench v0.1 | Claude Fable 5 | 33.7% | independent | test-lab.ai | test-lab.ai |
| Frontier-Bench v0.1 | Claude Opus 5 | 43.3% | independent | test-lab.ai | test-lab.ai |
| Terminal-Bench v2.1 | Grok 4.6 | 88.4% | independent | Artificial Analysis | artificialanalysis.ai |
| GDPval-AA v2 | Grok 4.6 | 1753 Elo | independent | Artificial Analysis | artificialanalysis.ai |
Method: evidence ledger of 777 filtered rows plus a 1,902-row merged ledger and a 6-model coverage summary, all harvested from public pages published to 2026-09-05. Matrix cells use only rows where the benchmark name and model-score association are explicit; modes (max, high, effort level) are preserved and unlike modes are flagged, not averaged. Provenance: I = independent measured evaluation, F = first-party benchmark claim repeated by an independent publication, A = anecdotal or secondary hands-on. Scores are percentages unless marked index or Elo. Blank cells mean no explicit row was found. Total evaluation-run cost is undisclosed corpus-wide (2 of 2,679 rows carry any cost) and left blank. Hundreds of rows without an explicit test name were cut for space.