No sweep: GPT-6 Astra takes the freshest overall boards by a nose, Claude Fable 5.1 owns agentic, coding and knowledge

Asked:“Who is winning Ai race at the moment chatgpt 6 or claude”

24 benchmark rows collected 10–24 September 2026 from three independent hosts pair OpenAI’s GPT-6 Astra against Anthropic’s Claude Fable 5.1 across eight head-to-head snapshots. Benchmarks test different capabilities on different scales, so each is shown on its own panel — these are snapshots, not a permanent championship.

OpenAI — GPT-6 AstraAnthropic — Claude Fable 5.1Each panel keeps its own scale and unit; leader named at top right
GPQA Diamond (science reasoning)GPT-6 Astra leads
% correct · 2026-09-24 · modelgrep.com
GPT-6 Astra rank 1
96.1%
Claude Fable 5.1 rank 8
93.7%
BenchAlign overall — September snapshotGPT-6 Astra leads
overall score · 2026-09-22 · benchlm.ai
GPT-6 Astra rank 1
88.47
Claude Fable 5.1 rank 3
83.34
Bug Hunt Bench — best variant per providerGPT-6 Astra leads
bugs fixed of 105 · 2026-09-10 · github.com
GPT-6 Astra rank 1
48 / 105
Fable 5.1 rank 3
43 / 105
BenchLM agenticClaude Fable 5.1 leads
weighted score · 2026-09-18 · benchlm.ai
Claude Fable 5.1 rank 1
80.2
GPT-6 Astra rank 5
70.3
BenchLM codingClaude Fable 5.1 leads
weighted score · 2026-09-18 · benchlm.ai
Claude Fable 5.1 rank 1
83.8
GPT-6 Astra rank 4
74.4
BenchLM knowledgeClaude Fable 5.1 leads
weighted score · 2026-09-18 · benchlm.ai
Claude Fable 5.1 rank 1
86.7
GPT-6 Astra rank 4
81.9
BenchAlign overall — August snapshotClaude Fable 5.1 leads
overall score · 2026-09-21 · benchlm.ai
Claude Fable 5.1 rank 1
84.58
GPT-6 Astra rank 2
83.79
BenchAlign overall — August snapshot (earlier read)Claude Fable 5.1 leads
overall score · 2026-09-10 · benchlm.ai
Claude Fable 5.1 rank 1
84.61
GPT-6 Astra rank 2
84.08
Freshest science snapshot: GPT-6 Astra tops GPQA Diamond at 96.1% while Fable 5.1 sits eighth at 93.7% — modelgrep.com
The overall board flipped in September: Fable 5.1 led the August BenchAlign reads (84.61 and 84.58 vs 84.08 and 83.79), then GPT-6 Astra jumped to 88.47 vs 83.34 — benchlm.ai
Workload splits are decisive for Anthropic: Fable 5.1 leads agentic by 9.9 points, coding by 9.4 and knowledge by 4.8 on BenchLM’s weighted boards — benchlm.ai
On Bug Hunt’s 105 real bugs, each provider ran several configurations; the best GPT-6 Astra run fixed 48, the best Fable 5.1 run 43 — github.com

All 24 rows

BenchmarkModelScoreRankDateSource
GPQA DiamondGPT-6 Astra96.1%12026-09-24modelgrep.com
GPQA DiamondClaude Fable 5.193.7%82026-09-24modelgrep.com
BenchAlign BenchLM leaderboardGPT-6 Astra88.4712026-09-22benchlm.ai
BenchAlign BenchLM leaderboardClaude Fable 5.183.3432026-09-22benchlm.ai
BenchAlign BenchLM leaderboardClaude Fable 5.184.5812026-09-21benchlm.ai
BenchAlign BenchLM leaderboardGPT-6 Astra83.7922026-09-21benchlm.ai
BenchLM's agentic leaderboardClaude Fable 5.180.2 weighted score12026-09-18benchlm.ai
BenchLM's agentic leaderboardGPT-6 Astra70.3 weighted score52026-09-18benchlm.ai
BenchLM's coding leaderboardClaude Fable 5.183.8 weighted score12026-09-18benchlm.ai
BenchLM's coding leaderboardGPT-6 Astra74.4 weighted score42026-09-18benchlm.ai
BenchLM's knowledge leaderboardClaude Fable 5.186.7 weighted score12026-09-18benchlm.ai
BenchLM's knowledge leaderboardGPT-6 Astra81.9 weighted score42026-09-18benchlm.ai
BenchAlign BenchLM leaderboardClaude Fable 5.184.6112026-09-10benchlm.ai
BenchAlign BenchLM leaderboardGPT-6 Astra84.0822026-09-10benchlm.ai
Bug Hunt BenchGPT-6 Astra48 /10512026-09-10github.com
Bug Hunt BenchGPT-6 Astra43 /10522026-09-10github.com
Bug Hunt BenchFable 5.143 /10532026-09-10github.com
Bug Hunt BenchGPT-6 Astra35 /10562026-09-10github.com
Bug Hunt BenchGPT-6 Astra34 /10572026-09-10github.com
Bug Hunt BenchFable 5.133 /105102026-09-10github.com
Bug Hunt BenchFable 5.129 /105132026-09-10github.com
Bug Hunt BenchFable 5.129 /105142026-09-10github.com
Bug Hunt BenchGPT-6 Astra27 /105162026-09-10github.com
Bug Hunt BenchFable 5.121 /105272026-09-10github.com

Method: 24 comparison rows extracted 2026-09-10 to 2026-09-24 from benchlm.ai (12 rows), github.com and modelgrep.com; each row is one model’s published score and rank on one benchmark snapshot. Scales differ by benchmark (percent correct, weighted 0–100 scores, bugs fixed of 105), so panels are not comparable across each other. Bug Hunt panels compare each provider’s best listed configuration; all ten Bug Hunt rows appear in the table. Nothing was cut.

This report was generated automatically by Keenable SELECT at a user's request, from publicly available web sources linked herein. Keenable does not review, verify, or endorse its contents and makes no representation as to accuracy, completeness, or timeliness; AI-based extraction may contain errors. Nothing in this report is investment, legal, financial, or other professional advice. All trademarks and referenced content remain the property of their respective owners; no affiliation or endorsement is implied. To report an error, rights concern, or request removal: legal@keenable.ai.

Keenable SELECTAsk your own question
Made with Keenable SELECT