Claude’s highest reported GPQA score rose 59.4 points in 27 months

Across nine major Claude releases with recoverable GPQA Diamond results, the reported high moved from 31.9% for Claude 2.1 to 91.3% for Opus 4.6. The biggest visible step came with Claude 3.7 Sonnet, though changing evaluation settings mean this is a history of published results—not a controlled rerun.

91.3%
latest selected high
+59.4
percentage-point gain
2.9×
score multiple
9
release points traced
Reported GPQA Diamond score by releasehighest recovered result, %
2040608010031.940.450.465.084.879.683.887.391.3
Claude 2.1
Nov ’23
3 Sonnet
Mar ’24
3 Opus
Mar ’24
3.5 Sonnet
Jun ’24
3.7 Sonnet
Feb ’25
Opus 4
May ’25
Sonnet 4
May ’25
Opus 4.5
Nov ’25
Opus 4.6
Feb ’26

What changed

1
The Claude 3 family established the first clear staircase. Sonnet’s 40.4% was followed by Opus at 50.4% under the reported GPQA Diamond results. langbase.com
2
Claude 3.7 Sonnet produced the largest milestone jump. Its highest recovered result was 84.8%, 19.8 points above the selected 3.5 Sonnet result. vellum.ai
3
Not every later model raised the line. Opus 4’s recovered 79.6% sat below 3.7 Sonnet’s high, while Sonnet 4 reached 83.8% with a deeper-thinking/tools setting. anthropic.com
4
Opus regained the frontier. Opus 4.5 reached 87.3%, followed by 91.3% for Opus 4.6 in the selected published result. claude5.ai

Three eras

baseline

Claude 2.1 → Claude 3

31.9 → 50.4

GPQA moved from low-thirties to roughly half correct by the first Opus generation.

reasoning jump

Claude 3.5 → 3.7

65.0 → 84.8

Reasoning modes drove the sharpest reported increase, but also reduced protocol comparability.

frontier consolidation

Opus 4.5 → 4.6

87.3 → 91.3

The selected Opus results pushed published performance above nine in ten questions.

Release detail

DateModelScoreReported settingSource
2023-11-21Claude 2.131.9%Not recoveredaiflashreport.com
2024-03-01Claude 3 Sonnet40.4%0-shot CoTlangbase.com
2024-03-04Claude 3 Opus50.4%0-shot CoT in family tablepromptlayer.com
2024-06-01Claude 3.5 Sonnet65.0%Standard modegetpassionfruit.com
2025-02-24Claude 3.7 Sonnet84.8%Highest of two reported settingsvellum.ai
2025-05-22Claude Opus 479.6%Highest recovered headline valueanthropic.com
2025-05-22Claude Sonnet 483.8%Deep thinking + toolsanthropic.com
2025-11-24Claude Opus 4.587.3%Not recoveredclaude5.ai
2026-02-05Claude Opus 4.691.3%Standard/no extended thinking in recovered sourceagentwiki.org

Method: web search run 25 Aug 2026 across Anthropic pages, benchmark aggregators and model-analysis sites; 1,052 unique search results were scanned and 30 normalized Claude-name candidates were assembled. Values are the highest recovered published GPQA Diamond result for each selected release, not a fresh uniform evaluation. Dates and protocols occasionally differ across sources, so compare the long-run direction more confidently than small adjacent gaps.

Keenable · made with SELECT* · 2,325 pages in 4m 19s · Open the chatShare: XLinkedInReddit