A web scan merged several thousand search hits into roughly 1,900 candidate model mentions across six result sets covering calendar years 2022–2026 (2026 is partial, through 3 September). The headline metric is word error rate (WER, %; lower is better) macro-averaged across the eight English test sets of the Hugging Face Open ASR Leaderboard — a macro-average WER is not literally “percent accuracy,” but lower WER means better recognition. Points are keyed to model release year, not the date of the page quoting them.
Best substantiated eight-set macro-average WER per release year, with other reported models per year shown as smaller dots. 2022 is a gap: releases like Whisper large-v2 predate the eight-set macro-average and no comparable 2022 figure appears in the rows.
One point per model and release year; filled bars are test-clean, hollow extensions are test-other WER. Decoding differs across points: E-Branchformer uses an external LM with ILME rescoring, and ECLM (2024, 1.5/3.3) is an error-correction pipeline applied after an ASR system, not a standalone acoustic model — single-benchmark rankings are therefore not a controlled longitudinal series.
Whisper set the open-model standard early in the window: Whisper large-v2 (Dec 2022) and Whisper Large v3 (7.44% eight-set average, 2023) were the reference points every later model was measured against.
Canary 1B (6.50%, 2024) and the Parakeet-TDT and Canary-Qwen families (6.05% and 5.63%, 2025) took the lead in 2024–2025 with compact FastConformer models, some under 1B parameters.
Cohere Transcribe leads in 2026 at 5.42%. Specialist and research contributors — Alibaba Qwen3-ASR, ElevenLabs, Mistral Voxtral, Microsoft Phi-4, IBM Granite, Samba-ASR — crowd within about one point of the frontier.
| Year | Model | Lab | Avg WER % | LS clean | LS other | Source |
|---|---|---|---|---|---|---|
| 2023 | Whisper Large v3 | OpenAI | 7.44 | — | — | codesota.com |
| 2024 | Canary 1B | NVIDIA | 6.50 | — | — | codesota.com |
| 2025 | Canary-Qwen-2.5B | NVIDIA | 5.63 | — | — | codesota.com |
| 2026 | Cohere Transcribe | Cohere | 5.42 | — | — | codesota.com |
| 2024 | Whisper Large v3 Turbo | OpenAI | 7.83 | — | — | codesota.com |
| 2024 | Parakeet TDT 1.1B | NVIDIA | 7.02 | — | — | codesota.com |
| 2025 | Parakeet TDT 0.6B v2 | NVIDIA | 6.05 | — | — | codesota.com |
| 2025 | Qwen3-ASR-1.7B | Alibaba | 5.76 | — | — | codesota.com |
| 2025 | ElevenLabs Scribe v2 | ElevenLabs | 5.83 | — | — | codesota.com |
| 2025 | Voxtral Small 24B | Mistral AI | 6.62 | — | — | codesota.com |
| 2025 | Phi-4 Multimodal | Microsoft | 6.02 | — | — | codesota.com |
| 2022 | Whisper large-v2 | OpenAI | — | 2.7 | 5.2 | arxiv.org |
| 2023 | E-Branchformer | Academic (SLT 2022, pub. Jan 2023) | — | 1.81 | 3.65 | ar5iv.labs.arxiv.org |
| 2024 | ECLM (error-correction pipeline) | Research | — | 1.5 | 3.3 | arxiv.org |
| 2025 | Samba-ASR | Research (Mamba SSM) | — | 1.17 | 2.48 | alphaxiv.org |
| 2025 | Parakeet TDT 0.6B v2 | NVIDIA | — | 1.69 | 3.19 | hearsy.app |
| 2025 | Canary-Qwen-2.5B | NVIDIA | — | 1.6 | 3.1 | gladia.io |
| 2026 | Cohere Transcribe | Cohere | — | 1.25 | 2.37 | marktechpost.com |
Method: a web scan merged several thousand search hits into six candidate result sets (~1,900 extracted model mentions after unnesting); null and implausible rows were filtered and repeated mentions deduplicated. Numbers are word error rate in percent (lower is better); the lead chart uses the eight-set macro-average of the Hugging Face Open ASR Leaderboard, the second chart LibriSpeech test-clean/test-other. Points are keyed to extracted original release year, corroborated against evidence text and URLs; newer pages quoting old models are not treated as new releases. 2022 has no substantiated eight-set macro-average and is shown as a gap; 2026 is partial-year data through 2026-09-03. Dozens of near-frontier models were cut for space.