Thirteen open-weight models over 30B now ship permissively with a reported GPQA score

Every open-weight model variant with more than 30 billion total parameters, an Apache 2.0, MIT, or Modified MIT license, weights public between 2026-02-01 and 2026-08-29, and a reported numeric GPQA Diamond score. Scores are vendor-reported unless noted; the >30B threshold uses total parameters, not active.

Apache 2.0 MIT Modified MIT — permits use with an attribution condition on code and weights Axis starts at 82, not zero, to separate a tight field
Kimi K3 leads at 93.5, but its weights only became public on July 27 — eleven days after the July 16 launch, so the weight-release date is the qualifying one. thundercompute.com marketintelligenceresearch.com
GLM-5.2's 91.2 is a self-reported figure; independent re-runs by Artificial Analysis land around 89 owing to different scaffolding — and its weights went public June 16, three days after the June 13 service launch. noze.it
Gemma 4 31B is the only dense model on the list — every other qualifier is a mixture-of-experts. Its official model-card figure is 84.3; an independent Artificial Analysis reasoning run reported 85.7. huggingface.co github.com
Not included for insufficient evidence or failed criteria: GLM-5 (Feb 12, MIT, 744B — no numeric GPQA score in the evidence), Qwen3.5-27B at 27B and Gemma 4 26B A4B at 25.2B total (below the 30B threshold), and Llama 4 (use-restricted Community license). z.ai huggingface.co

The qualifying shortlist

Model / variant Total params (active) Weights public License GPQA, reported Source
Kimi K3 (Moonshot AI)2.8T (~50B)2026-07-27 ³Modified MIT ¹93.5 · GPQA Diamondthundercompute.com
GLM-5.2 (Z.ai)~753B (~40B)2026-06-16 ²MIT91.2 · GPQA-Diamond, self-reported (~89 independent)noze.it
Kimi K2.6 (Moonshot AI)1T (32B)2026-04-20Modified MIT ¹90.5 · GPQA Diamondchatforest.com
Hy3 (Tencent)295B (21B)2026-07-06Apache 2.090.4 · GPQA Diamond, vendor-measuredai-beat.github.io
DeepSeek V4 Pro1.6T (49B)2026-04-24MIT90.1 · GPQA Diamonddemandsphere.com
Inkling-Small (Thinking Machines)276B (12B)2026-07-30Apache 2.089.5% · GPQA Diamond, effort 0.99, lab tableexplainx.ai
Qwen3.5-397B-A17B (Alibaba)397B (17B)2026-02-16Apache 2.088.4 · GPQA Diamondertas.ai · intelligibberish.com
DeepSeek V4 Flash284B (13B)2026-04-24MIT88.1% · GPQA Diamonddemandsphere.com
Inkling (Thinking Machines)975B (41B)2026-07-15Apache 2.087.2% · GPQA Diamond, effort 0.99, official model cardthinkingmachines.ai
Qwen3.5-122B-A10B (Alibaba)122B (10B)2026-02-24Apache 2.086.6 · GPQA Diamond, official table (one tracker lists 85.5)thesingularitypoint.substack.com
GLM-5.1 (Z.ai)744B (40B)2026-04-07MIT86.2 · GPQA-Diamond, model cardcodersera.com
Gemma 4 31B (Google DeepMind)30.7B dense2026-04-02Apache 2.084.3% · GPQA Diamond, official model card (85.7 independent reasoning run)infoq.com
Qwen3.5-35B-A3B (Alibaba)35B (3B)2026-02-24Apache 2.084.2 · GPQA Diamond, official model cardhuggingface.co

¹ Modified MIT. Not plain MIT: Moonshot AI's Modified MIT License covers both code and weights and adds an attribution condition beyond standard MIT (per the chatforest.com licensing section). It is included here as permissive but labelled distinctly.

² GLM-5.2 date. The model launched on Z.ai's GLM Coding Plan service on June 13, 2026; downloadable open weights followed on Hugging Face on June 16, 2026 — the weight-publication date is used for the cutoff.

³ Kimi K3 date. Launched July 16, 2026; full weights followed on July 27, 2026 under Modified MIT — the weight-publication date is used.

Methodology. Open-weight does not mean open-source: it means downloadable weights, and the license may still restrict use. Permissive here means standard Apache 2.0 or MIT; Modified MIT is admitted with the label above; Llama Community and other use-restricted licenses are excluded. The >30B threshold applies to total parameters, not active — so Gemma 4 31B (30.7B) qualifies while Qwen3.5-27B and Gemma 4 26B A4B (25.2B total) do not. "After January 2026" means weights public on or after 2026-02-01, up to the 2026-08-29 research date. Where GPQA figures disagree across sources, official model cards and repositories take precedence and the evaluation label is preserved; secondary or divergent figures are noted in the score cell. Qwen3.5-397B-A17B's Apache 2.0 license is evidenced at family level in a comparison table on aiproductivity.ai rather than a per-variant model card.

Data: 160 evidence rows extracted from official model cards, repositories, technical reports, launch posts, and secondary trackers across several thousand URL-deduplicated search results, research date 2026-08-29. Each score is a reported GPQA or GPQA Diamond figure (percent correct on graduate-level science questions), vendor-reported unless labelled otherwise. 13 model variants qualified after deduplicating repeated sources and naming variants; sub-threshold variants, restricted licenses, and candidates missing any of the four required attributes (notably GLM-5) were cut from the shortlist and listed in the exclusions note.

Keenable · made with SELECT · 4,787 pages in 6m 57s · Open the conversationShare: XLinkedInReddit