Permissively Licensed Open-Weight Models Over 30B With GPQA Scores

Eight qualifying checkpoints released between 1 February and 28 August 2026, each with more than 30B total parameters, weights under Apache‑2.0, MIT, or Modified MIT, and a reported GPQA result. Sizes run from a 31B dense model to a 1.6T mixture-of-experts; two models carry openly conflicting scores across sources.

Apache-2.0 weights MIT weights Modified MIT — extra terms apply license unresolved, not qualifying bubble area ≈ total parameters · pulsing = conflicting reports

What the picture cannot say alone

GLM-5.2 posts the best qualifying score, but sources disagree on its facts: the strongest exact evidence gives 744B total and 91.2% GPQA-Diamond on 16 June (techjacksolutions.com), while a secondary roundup lists ~753B and GPQA 88.5 (local-ai-zone.github.io).
DeepSeek V4-Pro's GPQA-Diamond spans a 14-point gap across pages — 90.1 in one comparison (thundercompute.com) against 76.1 in a launch write-up (tokenscost.com); the report shows both rather than choosing.
Efficiency, not size, sets the pace: DeepSeek V4-Flash reaches 88.1 with only 13B active of 284B total (thundercompute.com), and Inkling-Small hits 89% with 12B active of 276B under plain Apache-2.0 (felloai.com).
The only dense qualifier is Gemma 4 31B at 84.3% GPQA Diamond — the smallest model on the page, and the top of the single-GPU class per its reviewers (aiwiki.ai, digitalstrategy-ai.com).

The qualifying rows

Click a column header to sort. "Gemma 4 31B" and "Gemma 4-31B IT" are spelling variants of the same flagship checkpoint and are merged; separately released variants (V4-Pro vs V4-Flash, GLM-5 vs GLM-5.2) stay separate.

Model Released Total params Active License GPQA (variant) Source
GLM-52026-02-11744B40B MIT86.0 · GPQA-Diamond deepseekai.guide
DeepSeek V42026-03-031T32B Apache-2.079.3 · GPQA kersai.com
Gemma 4 31B2026-04-0231B dense31B Apache-2.084.3 · GPQA Diamond aiwiki.ai · deepresearch.ninja
Kimi K2.62026-04-201T~32B Modified MIT90.5 · GPQA Diamond tech-insider.org
DeepSeek V4-Pro2026-04-241.6T49B MIT 90.1 · GPQA Diamond
⚠ conflicting report: 76.1 GPQA-Diamond
thundercompute.com · tokenscost.com
DeepSeek V4-Flash2026-04-24284B13B MIT88.1 · GPQA Diamond thundercompute.com
GLM-5.22026-06-16744B
one secondary source: ~753B
40B MIT 91.2 · GPQA-Diamond
⚠ secondary source: 88.5
techjacksolutions.com · local-ai-zone.github.io
Inkling-Small2026-07-30276B12B Apache-2.089 · GPQA Diamond felloai.com
Unresolved — not in the qualifying table Kimi K3 (released 2026-07-16, 2.8T total / ~50B active, GPQA Diamond 93.5 per thundercompute.com) would top this list, but the evidence never establishes a permissive weights license; a hands-on report even flags "license fine print" without naming terms (runaihome.com). It appears on the chart as a dashed ghost. Kimi K2.5 is excluded because its extracted date is only "2026", which does not establish a post-January release.

Method: live-web scan as of 2026-08-28 across three result sets (a 3-row strict shortlist where all fields were extracted together, a 51-page structured evidence set, and 1 corroborating page on the GLM-5.2 discrepancy). Filter: total parameters strictly above 30B, weights under Apache-2.0, MIT, or Modified MIT, release date 2026-02-01 through 2026-08-28, and a reported GPQA score. GPQA values are as reported by each page — mostly the GPQA-Diamond variant, self-reported or via third-party evaluation, not independently reproduced here; benchmark settings may differ between sources. Spelling-only duplicates were merged; pages with no extractable model fields, models at or below 30B, and rows missing any of the four attributes were cut. Conflicting figures are shown, not averaged.

Keenable · made with SELECT · 4,119 pages in 6m 56s · Open the conversationShare: XLinkedInReddit