No single winner: Arcee is the small lab to watch, while Moonshot, DeepSeek and Z.ai own the open frontier
Asked:
“which small/startup/upcoming AI lab is building the best/most interesting LLMs/models nowadays?”
Snapshot as of 30 September 2026: a 10-row research shortlist (nine distinct labs — Z.ai/Zhipu appears twice and is merged here) plus 120 supporting rows of model detail, benchmark evidence and flagship claims across roughly a dozen emerging labs. The shortlist's ten rows span eight independent hosts; no single host backs more than two rows. Rankings in the shortlist reflect search relevance, not a league table.
Where each lab sits: demonstrated quality vs technical distinctiveness
Frontier challenger — heavily funded, near closed-frontier qualityGenuinely small lab — the distinctive betsEstablished scale-up
Positions are editorial and qualitative, placed from the evidence in the rows — not numeric scores. Hover or tap a lab for its standout model and strongest cited claim.
What the matrix cannot say
DeepSeek V4 Pro is #1 among open-weight models on the GDPval-AA agentic leaderboard at 1554 Elo, with a vendor-reported 80.6% SWE-Bench Verified and $0.435 per 1M input tokens on the efficiency frontier — codersera.com, gradually.ai
GLM 5.2's 51 on the Artificial Analysis Intelligence Index v4.1 (fifth overall) is independently run, not self-reported — the strongest third-party number in the whole shortlist — o-mega.ai
Arcee's Trinity Large is a fully custom design — 398B total, ~13B active, 4-of-256 routing, its own 200K-token vocabulary — but its headline agentic scores (94.7% τ²-Bench, 98.2% LiveCodeBench) are vendor-reported — insiderllm.com
Sakana's Fugu Max is not a monolithic LLM at all: a learned coordinator routes each request across a pool of open and specialized models at a flat $2/$6 per 1M tokens — API only, and not available in the EU or EEA — eesel.ai
The recommended labs, one card each
The research shortlist, row by row
Lab
Model
Cited claim
Availability
Source
Method: 130 collected rows as of 2026-09-30 — a 10-row research shortlist (one cited claim per row; ten rows over eight hosts), 40 rows on small/emerging labs, 40 comparative benchmark rows, and 40 flagship-evidence rows. Matrix positions are qualitative editorial placements, not measured scores; benchmark figures are labelled vendor-reported or independently run per their source. Duplicate Z.ai/Zhipu rows are merged and repeated Trinity-Large Hugging Face mirrors deduplicated. Model and benchmark names move fast — treat versions as a snapshot. Cut for space: repeated mirrors and legacy-model rows.