Open TTS Leaderboard: Scalable Evaluation for Multilingual Text-to-Speech and Voice Cloning
- ID
- 30350
- Status
- summarized
- Published
- 30 Sep 2026, 8:00 AM
- Fetched
- 01 Oct 2026, 12:09 AM
- Provider
- Hugging Face Blog
- Category
- developer-ai
- Original URL
- https://huggingface.co/blog/open-tts-leaderboard
- Source URL
- https://huggingface.co/blog/feed.xml
Summary
- Score
- 6.5
- Created
- 01 Oct 2026, 12:10 AM
- Tags
- Audience
- developersai_ml_learnersai_agent_usersvibe_coders
What happened
Hugging Face published the Open TTS Leaderboard, an objective-metric alternative to arena-style TTS rankings (TTS Arena v2, Artificial Analysis Voice Arena), which rank models by human pairwise votes and Elo/Bradley-Terry scoring. It scores models on intelligibility (WER/CER against a Qwen3 ASR transcript), speed (RTFx for batched offline inference on an H200, and time-to-first-audio for streaming batch size 1 on H200 and CPU), and speaker similarity (cosine similarity between WavLM speaker embeddings of generated audio and the reference clip). The stated motivation: over 8K TTS models sit on the Hugging Face Hub as of Sep 30, 2026, yet only 16 of 92 models on Artificial Analysis are open-weights, and vote-based evaluation takes weeks while objective metrics take hours.
Why it matters
If you are choosing a TTS model for a voice feature or agent, this gives you per-model WER/CER, RTFx, TTFA, and speaker-similarity numbers you can filter on instead of arena Elo - and TTFA is the number that actually decides whether a streaming voice agent feels laggy, since it is measured separately for batch size 1 on CPU and H200. The open-weights skew is the practical warning: arenas over-represent API models, so an open model you can self-host may be missing or mis-ranked there. Note the excerpt does not include any actual model scores, so treat this as a methodology and a place to look rather than a result.
Discussion angle
Objective metrics vs human preference for TTS: WER/CER and TTFA are reproducible and cheap, but they cannot tell you whether a voice sounds natural or right for your brand - so which metrics would you actually gate a model release on, and does the CPU TTFA number match what your users would feel?