“Which open-weight AI models have a permissive license, more than 30B parameters, a release date after January 2026, and a reported GPQA score? List all four attributes for each model.”
A live-web scan on 8 September 2026 found 24 cross-checked rows, each backed by at least two independent hosts; merging three family or size aliases leaves 21 canonical model releases from 11 February through 28 August 2026. Because no authoritative registry of open-weight releases exists, “all” means all qualifying releases this broad scan found. GPQA scores are on the Diamond variant unless noted, in percent; parameters are totals (active counts shown for MoE models).
Bubble area scales with log of total parameters (31B to 2,800B). Colour = weights license. Hover a bubble for full details.
| Model | Released | Parameters | License | GPQA | Hosts | Source |
|---|
Sources: cross-checked rows each supported by two or more URLs on independent hosts; the linked page is one representative source. GPQA is the Diamond variant throughout; DeepSeek-V4-Flash is vendor-reported at maximum reasoning effort.
Found in the broader scan but not independently corroborated, or excluded because the license wording is not a specific permissive license. Exact-30B models are excluded because the threshold is strictly more than 30B.
Third-party quantizations and checkpoint refreshes with their own reported score; not counted as separate underlying models.
Method: live-web scan on 2026-09-08 across official model cards and repositories, announcements, benchmark trackers, and independent coverage. Main table: 24 cross-checked rows (each on 2+ independent hosts) merged to 21 canonical models; appendices hold 9 single-source or license-unresolved candidates and 2 checkpoint variants from a 42-row broader scan. Eligibility uses total parameters, not active MoE parameters; “after January 2026” means strictly later than 2026-01-31; Apache-2.0, MIT, Modified MIT and OpenMDW-1.1 count as permissive. GPQA numbers are percent scores on the stated variant (Diamond throughout) and are not strictly comparable across differing inference budgets. No authoritative registry of open-weight releases exists, so coverage is what the scan found; no representative URL backs more than half the rows.