Eight qualifying checkpoints released between 1 February and 28 August 2026, each with more than 30B total parameters, weights under Apache‑2.0, MIT, or Modified MIT, and a reported GPQA result. Sizes run from a 31B dense model to a 1.6T mixture-of-experts; two models carry openly conflicting scores across sources.
Click a column header to sort. "Gemma 4 31B" and "Gemma 4-31B IT" are spelling variants of the same flagship checkpoint and are merged; separately released variants (V4-Pro vs V4-Flash, GLM-5 vs GLM-5.2) stay separate.
| Model | Released | Total params | Active | License | GPQA (variant) | Source |
|---|---|---|---|---|---|---|
| GLM-5 | 2026-02-11 | 744B | 40B | MIT | 86.0 · GPQA-Diamond | deepseekai.guide |
| DeepSeek V4 | 2026-03-03 | 1T | 32B | Apache-2.0 | 79.3 · GPQA | kersai.com |
| Gemma 4 31B | 2026-04-02 | 31B dense | 31B | Apache-2.0 | 84.3 · GPQA Diamond | aiwiki.ai · deepresearch.ninja |
| Kimi K2.6 | 2026-04-20 | 1T | ~32B | Modified MIT | 90.5 · GPQA Diamond | tech-insider.org |
| DeepSeek V4-Pro | 2026-04-24 | 1.6T | 49B | MIT | 90.1 · GPQA Diamond ⚠ conflicting report: 76.1 GPQA-Diamond |
thundercompute.com · tokenscost.com |
| DeepSeek V4-Flash | 2026-04-24 | 284B | 13B | MIT | 88.1 · GPQA Diamond | thundercompute.com |
| GLM-5.2 | 2026-06-16 | 744B one secondary source: ~753B | 40B | MIT | 91.2 · GPQA-Diamond ⚠ secondary source: 88.5 |
techjacksolutions.com · local-ai-zone.github.io |
| Inkling-Small | 2026-07-30 | 276B | 12B | Apache-2.0 | 89 · GPQA Diamond | felloai.com |
Method: live-web scan as of 2026-08-28 across three result sets (a 3-row strict shortlist where all fields were extracted together, a 51-page structured evidence set, and 1 corroborating page on the GLM-5.2 discrepancy). Filter: total parameters strictly above 30B, weights under Apache-2.0, MIT, or Modified MIT, release date 2026-02-01 through 2026-08-28, and a reported GPQA score. GPQA values are as reported by each page — mostly the GPQA-Diamond variant, self-reported or via third-party evaluation, not independently reproduced here; benchmark settings may differ between sources. Spelling-only duplicates were merged; pages with no extractable model fields, models at or below 30B, and rows missing any of the four attributes were cut. Conflicting figures are shown, not averaged.