21 permissively licensed open-weight models over 30B released after January 2026 report GPQA scores — Kimi K3 leads at 93.5

Asked:

“Which open-weight AI models have a permissive license, more than 30B parameters, a release date after January 2026, and a reported GPQA score? List all four attributes for each model.”

A live-web scan on 8 September 2026 found 24 cross-checked rows, each backed by at least two independent hosts; merging three family or size aliases leaves 21 canonical model releases from 11 February through 28 August 2026. Because no authoritative registry of open-weight releases exists, “all” means all qualifying releases this broad scan found. GPQA scores are on the Diamond variant unless noted, in percent; parameters are totals (active counts shown for MoE models).

Release date vs reported GPQA Diamond score

Bubble area scales with log of total parameters (31B to 2,800B). Colour = weights license. Hover a bubble for full details.

Kimi K3 is both the largest (2,800B total, 104B active) and the highest-scoring model at 93.5 GPQA Diamond, under a Modified MIT license — ai-tldr.dev
Scores climb through the year: every release after mid-June except Nemotron 3.5 Lightning and Motif 3 reports 87 or higher, versus 84–88 for the February wave — 7minai.com
Size is not destiny: Nex-N2-Pro reaches 90.7 with 397B total and only 17B active, beating 1.6T DeepSeek-V4-Pro — felloai.com
The two OpenMDW-1.1 Nemotron models are the clear outliers at 71.9 and 75.4, well below every MIT- and Apache-licensed peer — huggingface.co

All 21 canonical models — click a header to sort

ModelReleasedParametersLicenseGPQAHostsSource

Sources: cross-checked rows each supported by two or more URLs on independent hosts; the linked page is one representative source. GPQA is the Diamond variant throughout; DeepSeek-V4-Flash is vendor-reported at maximum reasoning effort.

Appendix: single-source or unresolved candidates

Found in the broader scan but not independently corroborated, or excluded because the license wording is not a specific permissive license. Exact-30B models are excluded because the threshold is strictly more than 30B.

Checkpoint appendix: quantizations and refreshes

Third-party quantizations and checkpoint refreshes with their own reported score; not counted as separate underlying models.

Method: live-web scan on 2026-09-08 across official model cards and repositories, announcements, benchmark trackers, and independent coverage. Main table: 24 cross-checked rows (each on 2+ independent hosts) merged to 21 canonical models; appendices hold 9 single-source or license-unresolved candidates and 2 checkpoint variants from a 42-row broader scan. Eligibility uses total parameters, not active MoE parameters; “after January 2026” means strictly later than 2026-01-31; Apache-2.0, MIT, Modified MIT and OpenMDW-1.1 count as permissive. GPQA numbers are percent scores on the stated variant (Diamond throughout) and are not strictly comparable across differing inference budgets. No authoritative registry of open-weight releases exists, so coverage is what the scan found; no representative URL backs more than half the rows.

This report was generated automatically by Keenable SELECT at a user's request, from publicly available web sources linked herein. Keenable does not review, verify, or endorse its contents and makes no representation as to accuracy, completeness, or timeliness; AI-based extraction may contain errors. Nothing in this report is investment, legal, financial, or other professional advice. All trademarks and referenced content remain the property of their respective owners; no affiliation or endorsement is implied. To report an error, rights concern, or request removal: legal@keenable.ai.

Keenable SELECTAsk your own question
Made with Keenable SELECT