Knowledge scores flattened near 90% while both labs made their biggest jumps on reasoning and coding
19 published benchmark snapshots for Anthropic and OpenAI models, November 2022 to May 2025 — MMLU knowledge, GPQA Diamond science reasoning, and SWE-bench coding, in percent. Evaluation setups differ across sources, so read the lines as directional history rather than one controlled experiment.
Anthropic
OpenAI
Hover any point for the exact benchmark, setting, and source
MMLU knowledge — first-to-last change
+18.7 pts
OpenAI · GPT-3.5 Turbo → GPT-4o · 70.0% → 88.7%
+10.5 pts
Anthropic · Claude 2 → Claude 3.7 Sonnet · 78.5% → 89.0%
GPQA Diamond reasoning — first-to-last change
+34.1 pts
OpenAI · GPT-4o → o3 · 53.6% → 87.7%
+26.0 pts
Anthropic · Claude 3 Opus → Claude Opus 4 · 48.9% → 74.9%
SWE-bench coding — first-to-last change
+39.4 pts
OpenAI · GPT-4o → o3 · 19.0% → 58.4%
+39.1 pts
Anthropic · Claude 3.5 Sonnet → Claude Opus 4 · 33.4% → 72.5%
What the timelines cannot say alone
MMLU is saturating: after GPT-4's 86.4% in March 2023, four further model generations across both labs added less than three points, all landing in the high 80s — a strong argument for the newer, harder benchmarks.
arxiv.org
OpenAI's reasoning-model turn produced the single biggest leap in the set: o3 reached 87.7% on GPQA Diamond, up 34.1 points from GPT-4o in under a year of releases.
aibytes.blog
Anthropic led every coding snapshot in this selection, climbing from 33.4% to 72.5% resolved across four releases, with Claude Opus 4's 72.5% the highest coding score in the data.
anthropic.com
Comparability caveat: OpenAI's coding line mixes SWE-bench (GPT-4o, 19.0% pass@1 in its system card) and SWE-bench Verified (o3, 58.4%) — related but not identical test sets, so the +39.4-point gain is directional.
openai.com
All 19 observations
| Benchmark | Model | Released | Score | Setting & source quality | Source |
Method: 19 published benchmark observations for Anthropic and OpenAI models (Nov 2022 – May 2025) plus 6 precomputed first-to-last deltas, assembled from official announcements and system cards where available and from benchmark papers and secondary compilations otherwise. Scores are accuracy percentages, or percent of issues resolved for SWE-bench. Dates are public release or announcement dates. Prompting, inference compute, tools, and scaffolds differ across sources; SWE-bench and SWE-bench Verified are related but not identical. Researched 2026-08-28; the timeline deliberately stops at well-supported 2025 observations and excludes poorly verified later claims.