Twenty-one post-January releases pair permissive weights with a reported GPQA score

A broad web scan found 21 distinct models meeting the strict filters after excluding Kimi K2.5, whose primary model card dates its release to January 29. Reported GPQA results range from 67.9 to 91.2, but differing variants and evaluation setups mean the scores are not a controlled leaderboard.

21
qualifying models
91.2
highest reported GPQA
14
Apache 2.0 releases
7
MIT releases

Reported GPQA score

Points or percent; GPQA Diamond unless source labels only GPQA
blue = vendor or cited reported score
GLM-5.2
91.2
Nex-N2-Pro
90.7
Tencent Hy3
90.4
MiMo V2.5-Pro
90.2
DeepSeek V4 Pro
90.1
Darwin-36B-Opus
90.0
Inkling Small
89.5
Qwen 3.5 397B
88.4
DeepSeek V4 Flash
88.1
Inkling
87.2
Qwen 3.5 122B
86.6
GLM-5.1
86.2
Qwen 3.6 35B
86.0
Gemma 4 31B
84.3
Qwen 3.5 35B
84.2
Motif 3
83.4
K-EXAONE 2.0
82.2
GLM-5
~78
Command A+
~76
Mistral Small 4
70.0
QwQ-32B-v2
67.9
0255075100

What stands out

1
GLM-5.2 has the highest reported result. Its MIT-licensed 753B MoE reports 91.2 on GPQA-Diamond. huggingface.co
2
Apache 2.0 is the dominant permissive license. Fourteen of the 21 qualifying models use Apache 2.0; seven use MIT.
3
Most are sparse MoEs. Many have hundreds of billions of total parameters but activate only 3B–49B per token, so total size is not inference cost.
4
The cutoff matters. Kimi K2.5 was excluded because its own model-card changelog gives January 29, 2026, despite later pages assigning it February or April dates. huggingface.co

Qualifying models

ModelReleaseParametersLicenseReported GPQA
Qwen 3.5 122B-A10B2026-02-01122B / 10B activeApache 2.086.6 Diamond · source
Qwen 3.5 35B-A3B2026-02-0135B / 3B activeApache 2.084.2 Diamond · source
Qwen 3.5 397B-A17B2026-02-01397B / 17B activeApache 2.088.4 · source
GLM-52026-02-11744B / 40B activeMIT~78 Diamond · source
Mistral Small 42026-03-16119B / 6.5B activeApache 2.00.70 reported (70%) · source
QwQ-32B-v22026-03-1632BApache 2.067.9% · source
Gemma 4 31B2026-04-0231BApache 2.084.3 Diamond · source
GLM-5.12026-04-07754B / 40B activeMIT86.2 Diamond · source
Darwin-36B-Opus2026-04-0836BApache 2.090.0 Diamond · source
Qwen 3.6 35B-A3B2026-04-1635B / 3B activeApache 2.086.0 Diamond · source
DeepSeek V4 Flash2026-04-24284B / 13B activeMIT88.1 Diamond · source
DeepSeek V4 Pro2026-04-241.6T / 49B activeMIT90.1 Diamond · source
Command A+2026-05-20218B / 25B activeApache 2.0~76 Diamond · source
Nex-N2-Pro2026-06-02397B / 17B activeApache 2.090.7 Diamond · source
Xiaomi MiMo V2.5-Pro 1T-A42B2026-06-031T / 42B activeMIT90.2 Diamond · source
GLM-5.22026-06-13753B / 40B activeMIT91.2 Diamond · source
Tencent Hy3 295B-A21B2026-07-06295B / 21B activeApache 2.090.4 Diamond · source
Inkling2026-07-15975B / 41B activeApache 2.087.2 Diamond · source
Inkling Small2026-07-30276B / 12B activeApache 2.089.5% · source
K-EXAONE 2.02026-08-05750B / 37B activeApache 2.082.2 Diamond · source
Motif 32026-08-18314B / 13.2B activeMIT83.4 Diamond · source
Method: broad web search across 2,472 deduplicated pages, checked 2026-08-23. “Parameters” means total model parameters; active parameters are shown for MoE models. Included only dates later than 2026-01-31 and standard permissive Apache 2.0 or MIT licenses. Scores are reported figures, not normalized independent retests.
Keenable · made with SELECT* · 3,931 pages in 3m 31s · Open the chatShare: XLinkedInReddit