Research and fact-check a draft article, “Voice AI Speech Models Are Moving Beyond Accuracy” (slug: voice-ai-speech-models-beyond-accuracy), on how speech-to-text competition is shifting from accuracy alone to latency, speakers, real-world audio and specialization — verifying its benchmark numbers, release timeline, and vendor claims.
The evidence set holds 76 rows captured on 2026-09-21: the current Pipecat speech-to-text benchmark (22 unique models from 15 providers after deduplicating identical rows repeated across pipecat.ai and GitHub — not the article’s 24), plus release and specialization records; 100 further source-level rows cross-check six proposed late-August/September releases. The Pipecat benchmark snapshot itself is single-sourced to Pipecat’s official page and repository; the wider report spans many independent hosts. WER is in percent, latency in milliseconds unless stated in seconds.
Pooled semantic WER (%) vs median time to final segment (ms), measured after end of speech, on 1,000 real-world samples. Lower-left is better; ★ marks the six Pareto-frontier models computed from these points. Hover or tap any point for its exact values.
Green: date verified first-party. Grey: independent coverage only. Amber: date conflicts with the article. Open circle: no supporting row — do not claim all six dates are verified.
Average diarization error rate falls 10.4% versus Precision-2, from 16.02 to 14.35 DER, across 15 benchmark datasets; adds and tunes voice-activity, crosstalk and probability controls.
Vendor-published benchmark suite · pyannote.ai
Positioned for noisy, real-world dictation; the company reports word error rates falling from over 30% to roughly 5–10% in noisy conditions on its own data. No independent evaluation appears in the captured rows — attribute these figures to Wispr.
Self-reported dataset and results · kompozy.io
60 input languages and 29 spoken output languages, with average lag reported falling from 2.8 s to 2.3 s, plus real-time speaker diarization.
Vendor-reported latency · thecosmicmeta.com
xAI reports short-phrase WER falling from 20.6% to 6.8%, and adds speaker labelling and multi-channel features (up to 8 channels in coverage). Figures are vendor-reported, not independently benchmarked in the rows.
Self-reported gains · kie.ai
| Provider | Model (★ frontier) | Pooled semantic WER % | Median TTFS ms | Source |
|---|---|---|---|---|
| Meta | muse-voice-transcribe-1.0 ★ | 0.83 | 392 | pipecat.ai |
| Speechmatics | (unlabeled) | 1.07 | 495 | pipecat.ai |
| Azure | (unlabeled) | 1.18 | 1016 | pipecat.ai |
| AssemblyAI | universal-3-5-pro ★ | 1.22 | 282 | pipecat.ai |
| Cartesia | ink-2 | 1.25 | 299 | pipecat.ai |
| Soniox | stt-rt-v5 ★ | 1.27 | 260 | pipecat.ai |
| Soniox | stt-rt-v4 ★ | 1.29 | 249 | pipecat.ai |
| AssemblyAI | u3-rt-pro | 1.34 | 335 | pipecat.ai |
| Deepgram | nova-3-general ★ | 1.62 | 247 | pipecat.ai |
| AWS | (unlabeled) | 1.75 | 1136 | pipecat.ai |
| NVIDIA | Nemotron 3.0 ASR (en) ★ | 1.95 | 221 | pipecat.ai |
| gemini-3.5-transcribe-live | 2.24 | 458 | pipecat.ai | |
| Smallest AI | pulse | 2.37 | 398 | pipecat.ai |
| OpenAI | gpt-realtime-whisper | 2.73 | 740 | pipecat.ai |
| latest-long | 2.85 | 878 | pipecat.ai | |
| AssemblyAI | universal-streaming-english | 3.02 | 256 | pipecat.ai |
| OpenAI | gpt-4o-transcribe | 3.06 | 637 | pipecat.ai |
| ElevenLabs | scribe_v2_realtime | 3.12 | 281 | pipecat.ai |
| Gradium | default | 3.71 | 570 | pipecat.ai |
| Cartesia | ink-whisper | 4.36 | 266 | pipecat.ai |
| NVIDIA | Nemotron 3.5 ASR (multilingual) | 4.58 | 236 | pipecat.ai |
| Mistral | voxtral-mini-transcribe-realtime-2602 | 4.97 | 525 | pipecat.ai |
Fact-check of one draft article against 76 evidence rows (benchmarks, releases, specialization) and 100 source-level release records, captured 2026-09-21. Chart shows 22 unique model entries from Pipecat’s official page and GitHub repository after deduplication; the benchmark snapshot is single-sourced to Pipecat. WER = pooled semantic word error rate in percent on 1,000 real-world samples; latency = median time to final segment in ms after end of speech. Duplicate rows across mirrors and repeated press rewrites were cut for space; a superseded GitHub README value for Gradium (3.96%) was dropped in favor of the current 3.71%.