216 public voice-AI benchmarks, split across seven fronts — and nearly half of last year's activity landed in 2026 alone

Asked:
"Do the same search for all the voice AI-related benchmarks."

A broad web search as of 2026-09-02 surfaced 216 publicly discoverable named voice-AI benchmark, leaderboard, arena and challenge series — the 28 STT/ASR series from the prior report plus 188 more, spanning speech generation, voice agents, safety, speaker tasks, speech quality and audio language models. 127 series carry an identifiable public date; 89 are undated on the discovered source.

The constellation: every series, by last identified update and category
One mark per benchmark series · x-axis: last identified public update (dated series) · right band: undated series · colour: category · hover a mark for detail
Seven fragments of one problem
Series per category (category counts as published in the catalog)
Generation is the loudest front: 56 of 216 series evaluate TTS, voice cloning or conversion — led by live human-preference arenas such as the Hugging Face / TTS-AGI TTS Arena.
Of the 127 dated series, 96 were last updated in 2026 and 20 in 2025 — the field's evaluation layer is being rebuilt almost in real time; only 11 dated series predate 2025 (e.g. the VoicePrivacy Challenge series traces back years but was still active in August 2026).
Fragmentation is structural: no single suite covers the landscape — even meta-collections like Dynamic-SUPERB sit beside 30 separate leaderboards/arenas and 30 recurring challenges, each with its own metrics and entrants.
89 of 216 series (41%) expose no reliable update date on their discovered public source — a transparency gap that makes freshness hard to judge; an example is the community TidyVoice 2026 Challenge results page.
The catalog: all 216 series
BenchmarkTask · measures · scopeVendors / projects representedLast update · source

Method: "all" means all named public comparative voice-AI benchmarks, leaderboards, arenas, benchmark suites, recurring challenges/shared tasks, and notable reproducible public benchmark studies discoverable in a broad web search as of 2026-09-02 — 216 series across STT/ASR; TTS, speech generation, cloning and conversion; speech-to-speech, full-duplex, voice agents and spoken dialogue; spoken/audio language models; speaker verification and diarization; emotion, accent, paralinguistics, VAD and keyword detection; speech quality, enhancement, separation and codecs; and deepfake/spoof detection, watermarking, privacy and anonymization. Excluded: bare datasets without a benchmark protocol or results, generic metrics, single-model evaluations, untested product listicles, pricing-only pages, music-only or environmental-audio-only evaluations, and private tests. There is no universal registry, so web-wide completeness cannot be guaranteed; unindexed, private, renamed or removed benchmarks may be missing. Annual editions and obvious aliases were merged into series where possible. Dates are stated benchmark/repository dates when available, otherwise the latest relevant public page or report date found; "undated" means no reliable update date was exposed, not that the benchmark changed on the cutoff. Some source pages are papers or secondary pages where the primary leaderboard was not indexable. Evidence-page count is a discovery/corroboration signal, not a quality score, and benchmarks are not ranked by it. Blank fields are preserved as found.

This report was generated automatically by Keenable SELECT at a user's request, from publicly available web sources linked herein. Keenable does not review, verify, or endorse its contents and makes no representation as to accuracy, completeness, or timeliness; AI-based extraction may contain errors. Nothing in this report is investment, legal, financial, or other professional advice. All trademarks and referenced content remain the property of their respective owners; no affiliation or endorsement is implied. To report an error, rights concern, or request removal: legal@keenable.ai.

Keenable SELECTAsk your own question
Made with Keenable SELECT