In fifteen months, Claude went from fixing one in five GitHub issues to four in five — then the climb got hard
Nine Anthropic-reported results on SWE-bench Verified, the software-engineering benchmark measuring the share of real GitHub issues a model resolves. Report dates span 30 Oct 2024 to 5 Feb 2026; units are percent resolved. High-compute parallel-sampling results are excluded.
22.0% → 81.42%Claude 3 Opus to Claude Opus 4.6, percent of issues resolved
+59.42 ptstotal gain across the nine-point timeline
+270.1%relative improvement — 3.70× the starting score
steep phase — gains of 9 to 16 points per step
flattening phase — gains of 0.52 to 3.7 points per step
hover a point for score, gain, and the methodology behind it
Read this as a release history, not a controlled experiment. All scores are Anthropic-reported, and the evaluation details evolved across releases — agent scaffolds, thinking budgets, number of trials (a 10-trial average for Sonnet 4.5, a 25-trial average with a noted prompt modification for Opus 4.6), and prompt wording. The first three results were reported together retrospectively on 30 Oct 2024, so their shared report date is not each model's original release date.
The first four steps did most of the work: 22.0 → 33.0 → 49.0 → 63.7 added 41.7 of the 59.42 total points, with the single largest jump (+16.0) between the two Claude 3.5 Sonnet versions.
anthropic.com
Claude 3.7 Sonnet's 63.7% was achieved without the earlier custom scaffold, per Anthropic's release note — a methodology shift mid-timeline.
anthropic.com
The last five results — 72.7, 74.5, 77.2, 80.9, 81.42 — span nine months for +8.72 points combined, less than the smallest single early-phase jump.
anthropic.com
The final step is the smallest on record: Opus 4.5 → Opus 4.6 gained just 0.52 points, on a 25-trial average with a prompt modification noted by Anthropic.
anthropic.com
All nine results
| Model |
Reported |
Score % |
Gain, pts |
Methodology note |
Source |
Source: official Anthropic benchmark announcements and system cards, checked through 2026-08-28. 9 rows, one per selected Claude model result on SWE-bench Verified. Score is the percent of the benchmark's real GitHub issues resolved; gain is percentage points over the prior point in this timeline. Excluded for comparability: high-compute parallel-sampling results. The first three scores share a retrospective report date of 2024-10-30. Methodology notes are abridged in the table; full wording is in the linked sources.