ASR leadership passed from OpenAI to NVIDIA and then to Cohere, whose Transcribe holds the top spot at 5.42% average WER

Asked:
“Which labs have created the most accurate speech recognition models, and how has accuracy changed over the past 5 years year-to-year?”

A web scan merged several thousand search hits into roughly 1,900 candidate model mentions across six result sets covering calendar years 2022–2026 (2026 is partial, through 3 September). The headline metric is word error rate (WER, %; lower is better) macro-averaged across the eight English test sets of the Hugging Face Open ASR Leaderboard — a macro-average WER is not literally “percent accuracy,” but lower WER means better recognition. Points are keyed to model release year, not the date of the page quoting them.

OpenAINVIDIACohereother labs

Best substantiated eight-set macro-average WER per release year, with other reported models per year shown as smaller dots. 2022 is a gap: releases like Whisper large-v2 predate the eight-set macro-average and no comparable 2022 figure appears in the rows.

LibriSpeech test-clean and test-other tell a different story per benchmark

One point per model and release year; filled bars are test-clean, hollow extensions are test-other WER. Decoding differs across points: E-Branchformer uses an external LM with ILME rescoring, and ECLM (2024, 1.5/3.3) is an error-correction pipeline applied after an ASR system, not a standalone acoustic model — single-benchmark rankings are therefore not a controlled longitudinal series.

Cohere Transcribe (5.42%) cuts average errors about 10.4% relative to NVIDIA Parakeet-TDT-0.6B-v2 (6.05%) and about 16.6% relative to NVIDIA Canary 1B (6.50%). codesota.com
Improvement is flattening: the best macro-average fell 0.94 points from 2023 to 2024, 0.87 from 2024 to 2025, but only 0.21 from 2025 to 2026 YTD. codesota.com
Cohere Transcribe, a 2B Apache-2.0 model released 26 March 2026, took #1 on the Open ASR Leaderboard immediately at launch, with the best LibriSpeech scores on the board (1.25 clean / 2.37 other). byteiota.com
On LibriSpeech alone, research systems such as Samba-ASR (1.17/2.48, Jan 2025) beat every production model — a reminder that “most accurate” depends on benchmark, domain, streaming constraints, and whether external LMs or post-correction are allowed. alphaxiv.org

The labs at the frontier

OpenAI

Whisper set the open-model standard early in the window: Whisper large-v2 (Dec 2022) and Whisper Large v3 (7.44% eight-set average, 2023) were the reference points every later model was measured against.

NVIDIA

Canary 1B (6.50%, 2024) and the Parakeet-TDT and Canary-Qwen families (6.05% and 5.63%, 2025) took the lead in 2024–2025 with compact FastConformer models, some under 1B parameters.

Cohere & specialists

Cohere Transcribe leads in 2026 at 5.42%. Specialist and research contributors — Alibaba Qwen3-ASR, ElevenLabs, Mistral Voxtral, Microsoft Phi-4, IBM Granite, Samba-ASR — crowd within about one point of the frontier.

Evidence table

YearModelLabAvg WER %LS cleanLS otherSource
2023Whisper Large v3OpenAI7.44codesota.com
2024Canary 1BNVIDIA6.50codesota.com
2025Canary-Qwen-2.5BNVIDIA5.63codesota.com
2026Cohere TranscribeCohere5.42codesota.com
2024Whisper Large v3 TurboOpenAI7.83codesota.com
2024Parakeet TDT 1.1BNVIDIA7.02codesota.com
2025Parakeet TDT 0.6B v2NVIDIA6.05codesota.com
2025Qwen3-ASR-1.7BAlibaba5.76codesota.com
2025ElevenLabs Scribe v2ElevenLabs5.83codesota.com
2025Voxtral Small 24BMistral AI6.62codesota.com
2025Phi-4 MultimodalMicrosoft6.02codesota.com
2022Whisper large-v2OpenAI2.75.2arxiv.org
2023E-BranchformerAcademic (SLT 2022, pub. Jan 2023)1.813.65ar5iv.labs.arxiv.org
2024ECLM (error-correction pipeline)Research1.53.3arxiv.org
2025Samba-ASRResearch (Mamba SSM)1.172.48alphaxiv.org
2025Parakeet TDT 0.6B v2NVIDIA1.693.19hearsy.app
2025Canary-Qwen-2.5BNVIDIA1.63.1gladia.io
2026Cohere TranscribeCohere1.252.37marktechpost.com

Method: a web scan merged several thousand search hits into six candidate result sets (~1,900 extracted model mentions after unnesting); null and implausible rows were filtered and repeated mentions deduplicated. Numbers are word error rate in percent (lower is better); the lead chart uses the eight-set macro-average of the Hugging Face Open ASR Leaderboard, the second chart LibriSpeech test-clean/test-other. Points are keyed to extracted original release year, corroborated against evidence text and URLs; newer pages quoting old models are not treated as new releases. 2022 has no substantiated eight-set macro-average and is shown as a gap; 2026 is partial-year data through 2026-09-03. Dozens of near-frontier models were cut for space.

This report was generated automatically by Keenable SELECT at a user's request, from publicly available web sources linked herein. Keenable does not review, verify, or endorse its contents and makes no representation as to accuracy, completeness, or timeliness; AI-based extraction may contain errors. Nothing in this report is investment, legal, financial, or other professional advice. All trademarks and referenced content remain the property of their respective owners; no affiliation or endorsement is implied. To report an error, rights concern, or request removal: legal@keenable.ai.

Keenable SELECTAsk your own question
Made with Keenable SELECT