Nine released variants clear all four filters; Darwin leads reported GPQA at 90.0%
Primary and developer-controlled sources yielded nine model variants released after 31 January 2026 with more than 30B total parameters, an MIT, Apache 2.0, or OpenMDW 1.1 license, and a reported GPQA-family result. Treat scores as reported rather than directly comparable: evaluation settings and numerical precision differ.
9
qualifying variants
90.0%
highest reported GPQA
35B–1.6T
total parameter span
3
permissive license families
Reported GPQA scores cluster between 84 and 90
Blue = reported score (%)
Reported GPQA Diamond unless the source labels the result simply GPQA; no score normalization was applied.
What stands out
01
Darwin posts the highest number. Its model page reports 90.0% GPQA Diamond for a 35B Apache-2.0 release. huggingface.co
02
Scale does not determine the reported result. The 35B Darwin score exceeds LongCat-2.0’s 88.9 at 1.6T total parameters. github.com
03
Precision creates distinct benchmark rows. NVIDIA reports 87.0 for Nemotron BF16 and 87.9 for NVFP4, so both released artifacts remain visible. nvidia.com
Qualifying models and all four requested attributes
| Model | Released | Parameters | License | Reported GPQA | Source |
|---|---|---|---|---|---|
| GLM-5 | 2026-02-12 | 744B total; 40B active | MIT | 86.0, GPQA-Diamond | z.ai |
| Qwen3.5-397B-A17B-NVFP4 | 2026-02-17 | 397B total; 17B active | Apache 2.0 | 87.1, GPQA Diamond | huggingface.co |
| Qwen3.5-35B-A3B | 2026-02-25 | 35B total; 3B active | Apache 2.0 | 84.2, GPQA Diamond | huggingface.co |
| Darwin-35B-A3B-Opus | 2026-05-14 | 35B total; 3B active | Apache 2.0 | 90.0%, GPQA Diamond | huggingface.co |
| Nemotron 3 Ultra BF16 | 2026-06-04 | 550B total; 55B active | OpenMDW 1.1 | 87.0, GPQA, no tools | nvidia.com |
| Nemotron 3 Ultra NVFP4 | 2026-06-04 | 550B total; 55B active | OpenMDW 1.1 | 87.9, GPQA, no tools | nvidia.com |
| GLM-5.2-NVFP4 | 2026-06-25 | 753B total; 40B active | MIT | 89.39, GPQA Diamond | huggingface.co |
| LongCat-2.0 | 2026-06-29 | 1.6T total; ~48B active | MIT | 88.9, GPQA-Diamond | github.com |
| Gemma 4 31B Dense | 2026-07-02 | 30.7B total | Apache 2.0 | 84.3%, GPQA Diamond | ai.google.dev |
Method: searched 1,849 web results across official model cards, technical reports, repositories, and developer announcements on 21 August 2026; retained released artifacts dated after 31 January, strictly above 30B total parameters, under MIT, Apache 2.0, or OpenMDW 1.1, with a reported GPQA-family score. Quantized and precision-specific artifacts are separate when their reported scores differ.