24 benchmark rows collected 10–24 September 2026 from three independent hosts pair OpenAI’s GPT-6 Astra against Anthropic’s Claude Fable 5.1 across eight head-to-head snapshots. Benchmarks test different capabilities on different scales, so each is shown on its own panel — these are snapshots, not a permanent championship.
| Benchmark | Model | Score | Rank | Date | Source |
|---|---|---|---|---|---|
| GPQA Diamond | GPT-6 Astra | 96.1% | 1 | 2026-09-24 | modelgrep.com |
| GPQA Diamond | Claude Fable 5.1 | 93.7% | 8 | 2026-09-24 | modelgrep.com |
| BenchAlign BenchLM leaderboard | GPT-6 Astra | 88.47 | 1 | 2026-09-22 | benchlm.ai |
| BenchAlign BenchLM leaderboard | Claude Fable 5.1 | 83.34 | 3 | 2026-09-22 | benchlm.ai |
| BenchAlign BenchLM leaderboard | Claude Fable 5.1 | 84.58 | 1 | 2026-09-21 | benchlm.ai |
| BenchAlign BenchLM leaderboard | GPT-6 Astra | 83.79 | 2 | 2026-09-21 | benchlm.ai |
| BenchLM's agentic leaderboard | Claude Fable 5.1 | 80.2 weighted score | 1 | 2026-09-18 | benchlm.ai |
| BenchLM's agentic leaderboard | GPT-6 Astra | 70.3 weighted score | 5 | 2026-09-18 | benchlm.ai |
| BenchLM's coding leaderboard | Claude Fable 5.1 | 83.8 weighted score | 1 | 2026-09-18 | benchlm.ai |
| BenchLM's coding leaderboard | GPT-6 Astra | 74.4 weighted score | 4 | 2026-09-18 | benchlm.ai |
| BenchLM's knowledge leaderboard | Claude Fable 5.1 | 86.7 weighted score | 1 | 2026-09-18 | benchlm.ai |
| BenchLM's knowledge leaderboard | GPT-6 Astra | 81.9 weighted score | 4 | 2026-09-18 | benchlm.ai |
| BenchAlign BenchLM leaderboard | Claude Fable 5.1 | 84.61 | 1 | 2026-09-10 | benchlm.ai |
| BenchAlign BenchLM leaderboard | GPT-6 Astra | 84.08 | 2 | 2026-09-10 | benchlm.ai |
| Bug Hunt Bench | GPT-6 Astra | 48 /105 | 1 | 2026-09-10 | github.com |
| Bug Hunt Bench | GPT-6 Astra | 43 /105 | 2 | 2026-09-10 | github.com |
| Bug Hunt Bench | Fable 5.1 | 43 /105 | 3 | 2026-09-10 | github.com |
| Bug Hunt Bench | GPT-6 Astra | 35 /105 | 6 | 2026-09-10 | github.com |
| Bug Hunt Bench | GPT-6 Astra | 34 /105 | 7 | 2026-09-10 | github.com |
| Bug Hunt Bench | Fable 5.1 | 33 /105 | 10 | 2026-09-10 | github.com |
| Bug Hunt Bench | Fable 5.1 | 29 /105 | 13 | 2026-09-10 | github.com |
| Bug Hunt Bench | Fable 5.1 | 29 /105 | 14 | 2026-09-10 | github.com |
| Bug Hunt Bench | GPT-6 Astra | 27 /105 | 16 | 2026-09-10 | github.com |
| Bug Hunt Bench | Fable 5.1 | 21 /105 | 27 | 2026-09-10 | github.com |
Method: 24 comparison rows extracted 2026-09-10 to 2026-09-24 from benchlm.ai (12 rows), github.com and modelgrep.com; each row is one model’s published score and rank on one benchmark snapshot. Scales differ by benchmark (percent correct, weighted 0–100 scores, bugs fixed of 105), so panels are not comparable across each other. Bug Hunt panels compare each provider’s best listed configuration; all ten Bug Hunt rows appear in the table. Nothing was cut.