Claude 2.1 → Claude 3
GPQA moved from low-thirties to roughly half correct by the first Opus generation.
Across nine major Claude releases with recoverable GPQA Diamond results, the reported high moved from 31.9% for Claude 2.1 to 91.3% for Opus 4.6. The biggest visible step came with Claude 3.7 Sonnet, though changing evaluation settings mean this is a history of published results—not a controlled rerun.
GPQA moved from low-thirties to roughly half correct by the first Opus generation.
Reasoning modes drove the sharpest reported increase, but also reduced protocol comparability.
The selected Opus results pushed published performance above nine in ten questions.
| Date | Model | Score | Reported setting | Source |
|---|---|---|---|---|
| 2023-11-21 | Claude 2.1 | 31.9% | Not recovered | aiflashreport.com |
| 2024-03-01 | Claude 3 Sonnet | 40.4% | 0-shot CoT | langbase.com |
| 2024-03-04 | Claude 3 Opus | 50.4% | 0-shot CoT in family table | promptlayer.com |
| 2024-06-01 | Claude 3.5 Sonnet | 65.0% | Standard mode | getpassionfruit.com |
| 2025-02-24 | Claude 3.7 Sonnet | 84.8% | Highest of two reported settings | vellum.ai |
| 2025-05-22 | Claude Opus 4 | 79.6% | Highest recovered headline value | anthropic.com |
| 2025-05-22 | Claude Sonnet 4 | 83.8% | Deep thinking + tools | anthropic.com |
| 2025-11-24 | Claude Opus 4.5 | 87.3% | Not recovered | claude5.ai |
| 2026-02-05 | Claude Opus 4.6 | 91.3% | Standard/no extended thinking in recovered source | agentwiki.org |
Method: web search run 25 Aug 2026 across Anthropic pages, benchmark aggregators and model-analysis sites; 1,052 unique search results were scanned and 30 normalized Claude-name candidates were assembled. Values are the highest recovered published GPQA Diamond result for each selected release, not a fresh uniform evaluation. Dates and protocols occasionally differ across sources, so compare the long-run direction more confidently than small adjacent gaps.