AI Weekly Malaysia

Back to items Summaries

Open TTS Leaderboard: Scalable Evaluation for Multilingual Text-to-Speech and Voice Cloning

ID
30350
Status
summarized
Published
30 Sep 2026, 8:00 AM
Fetched
01 Oct 2026, 12:09 AM
Provider
Hugging Face Blog
Category
developer-ai
Original URL
https://huggingface.co/blog/open-tts-leaderboard
Source URL
https://huggingface.co/blog/feed.xml

Summary

Score
6.5
Created
01 Oct 2026, 12:10 AM
Tags
Audience
developersai_ml_learnersai_agent_usersvibe_coders

What happened

Hugging Face published the Open TTS Leaderboard, an objective-metric alternative to arena-style TTS rankings (TTS Arena v2, Artificial Analysis Voice Arena), which rank models by human pairwise votes and Elo/Bradley-Terry scoring. It scores models on intelligibility (WER/CER against a Qwen3 ASR transcript), speed (RTFx for batched offline inference on an H200, and time-to-first-audio for streaming batch size 1 on H200 and CPU), and speaker similarity (cosine similarity between WavLM speaker embeddings of generated audio and the reference clip). The stated motivation: over 8K TTS models sit on the Hugging Face Hub as of Sep 30, 2026, yet only 16 of 92 models on Artificial Analysis are open-weights, and vote-based evaluation takes weeks while objective metrics take hours.

Why it matters

If you are choosing a TTS model for a voice feature or agent, this gives you per-model WER/CER, RTFx, TTFA, and speaker-similarity numbers you can filter on instead of arena Elo - and TTFA is the number that actually decides whether a streaming voice agent feels laggy, since it is measured separately for batch size 1 on CPU and H200. The open-weights skew is the practical warning: arenas over-represent API models, so an open model you can self-host may be missing or mis-ranked there. Note the excerpt does not include any actual model scores, so treat this as a methodology and a place to look rather than a result.

Discussion angle

Objective metrics vs human preference for TTS: WER/CER and TTFA are reproducible and cheap, but they cannot tell you whether a voice sounds natural or right for your brand - so which metrics would you actually gate a model release on, and does the CPU TTFA number match what your users would feel?

Top