The Open TTS Leaderboard has been developed to provide a more scalable and standardized method for evaluating text-to-speech (TTS) and voice cloning models. Existing evaluation methods, primarily human preference scores from arena-based leaderboards, face challenges with scalability and consistent voter criteria. These arenas require users to compare model outputs and vote, which is time-consuming and limits the number of models that can be assessed.
Human preference leaderboards, such as Artificial Analysis and Voice Arena, are considered the gold standard but cannot keep pace with new TTS releases. As of September 30, 2026, only a small fraction of models on these platforms are open-source, partly because open models require hosting by the arena operator, unlike API models. Additionally, voter consistency is a significant issue, as individual preferences can change over time, making consistent evaluation difficult.
The Open TTS Leaderboard utilizes objective metrics to evaluate models, significantly reducing evaluation time from weeks to hours. It assesses intelligibility using word/character error rate (WER and CER) with Qwen3 ASR, speed via inverse real-time factor (RTFx) and time-to-first-audio (TTFA) on H200 GPUs and CPUs, and speaker similarity by computing cosine similarity between WavLM speaker embeddings.
While the Open TTS Leaderboard provides objective measurements for specific performance aspects, it does not replace human preference ranking entirely. Metrics like ASR-based WER proxy intelligibility, and speaker similarity estimates voice identity preservation, but neither directly measures naturalness. The objective metrics offer a complementary approach to quickly assess technical performance characteristics of TTS models.
✨ This summary was generated by AI from the outlets' reporting listed below. It is not independently verified and may contain errors — check the original sources. How BrevFeed works →
One email each morning: the day's tech stories, clustered across outlets and summarized. No account needed.
One email a day. Unsubscribe in one click, any time.
Spend a few minutes, get the whole day. Every topic's top stories in one hands-free rundown — listen, watch, or read the transcript.
▶ Play today's briefNew every morning, and the back catalogue is archived by date.
A new Open TTS Leaderboard has been introduced to provide scalable, objective evaluation for text-to-speech and voice cloning models. It uses metrics like word error rate, inference speed, and speaker similarity to address the limitations of human preference-based arena leaderboards, which struggle with scalability and open-source model representation.