🔍 Read the full analysis: What Is The Open TTS Leaderboard? A Look At Scalable TTS Evaluation on ThorstenMeyerAI.com
Get tech for your team delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
Hugging Face has introduced the Open TTS Leaderboard, which compares text-to-speech models using automated measures of speech accuracy, generation speed and speaker similarity. The project says those results can be produced in hours, but they do not establish which voice sounds most natural or which system listeners prefer.
Hugging Face has launched the Open TTS Leaderboard, a system for comparing text-to-speech models on speech accuracy, generation speed and speaker similarity. The project is intended to make evaluations faster and more repeatable in a field where Hugging Face said its Hub contained more than 8,000 TTS models as of September 30, 2026, but the scores do not replace human judgments of naturalness or listener preference.
The leaderboard uses several measurements rather than a single overall measure of voice quality. For speech accuracy, it compares transcripts generated from model audio with the original prompts, reporting word error rate and character error rate using Qwen3 automatic speech recognition. For offline generation, it reports inverse real-time factor on an H200 GPU. For streaming, it measures time to first audio on both an H200 GPU and a CPU. A separate voice-cloning measure compares generated speech with reference audio using WavLM embeddings.
By default, the rankings use macro-average word error rate across English splits from Seed TTS Eval and CV3 Eval. Users can select other languages and switch to a voice-cloning view. Hugging Face identifies k2-fsa/OmniVoice, fishaudio/s2-pro and FunAudioLLM/Fun-CosyVoice3-0.5B-2512 as strong multilingual models. For English error rates, it lists hexgrad/Kokoro-82M, Supertone/supertonic-3 and fishaudio/s2-pro among the leaders. These are results on the leaderboard’s selected measures, not a general ranking of which systems sound best.
The project says an evaluation using its automated metrics can take a couple of hours, compared with weeks for arena voting. Its “Listen” tab lets users compare generated audio and submit preferences. Hugging Face asks voters to sign in with an account to limit spam and bot submissions, and says community votes may be incorporated into rankings later.
Faster Comparisons for TTS Teams
A shared set of repeatable tests may help developers and researchers narrow the field before spending time listening to outputs or integrating a model. The measures cover distinct practical concerns: accuracy can indicate whether spoken text matches a prompt, while speed and time to first audio can help teams assess suitability for applications with latency constraints, including voice agents. Speaker similarity offers another reference point for models used in voice cloning.
The metrics also make tradeoffs easier to see, but do not settle them. A model with a favorable error rate may not be the fastest, and a speaker-similarity score does not directly establish that speech sounds natural, expressive or appropriate to a listener. The leaderboard is most useful as a screening and comparison tool, with listening tests still needed for judgments tied to human experience.
Hugging Face says existing arena-style comparisons can be slow to expand because they require votes and, in some cases, hosting the models being tested. It also says open-weight models are underrepresented in some rankings. If the new system brings more models and languages into comparable evaluations, it could give teams a broader starting point; the announced information does not yet show how complete that coverage will be.
text-to-speech voice cloning software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Automated Scores and Listener Arenas
Text-to-speech systems convert written prompts into audio, and the number of model releases makes consistent comparison difficult. Existing reference points cited by Hugging Face include TTS Arena v2, Artificial Analysis and Voice Arena. Arena-style rankings typically ask listeners to compare outputs, collect votes and derive ratings. Those comparisons capture preferences directly, but their results depend on participation and the criteria voters apply.
The Open TTS Leaderboard takes a different approach for its main scores: it applies automated measures to standardized evaluation data, while retaining a listening feature for people to compare audio. Hugging Face said that, as of September 30, 2026, it counted 16 open-weight models among 92 on Artificial Analysis, and reported a similar imbalance on Voice Arena. That count and the explanation for it are Hugging Face’s account of the existing rankings, not an independent audit presented with the leaderboard announcement.
“The Open TTS Leaderboard does not replace human preference ranking.”
— Hugging Face, describing the leaderboard’s role
As an affiliate, we earn on qualifying purchases.
Limits of the Published Rankings
The announcement does not provide the full model list, score tables, evaluation sample sizes or uncertainty ranges for individual results. It also does not establish how closely the automated measures track listener judgments across languages, accents and speaking styles. Error rates estimate intelligibility through an automatic speech-recognition system, while WavLM-based similarity estimates voice identity resemblance; neither directly measures naturalness or expressiveness.
Coverage differs by language. The leaderboard uses character error rate for Chinese, Japanese and Korean. Hugging Face says that for languages beyond English and Chinese, Seed TTS Eval has no audio and scores come from CV3 Eval alone. The announcement does not specify how often rankings will refresh, how model or dataset changes will be handled, or when community votes might affect results. Those details limit how confidently readers can interpret small score differences or compare rankings over time.
As an affiliate, we earn on qualifying purchases.
Listening Tests and Ranking Updates
Users can explore the “Listen” tab, select a language and dataset, choose whether to compare voice cloning, and vote on generated audio. Hugging Face says it may incorporate accumulated community votes into the leaderboard, but has not announced a timetable or a threshold for doing so. Account sign-in is requested to reduce spam and bot activity.
The next useful signals will be fuller score reporting, details about evaluation coverage and a stated refresh process. Until those are available, the leaderboard’s automated results can help identify models for closer testing, while listening remains necessary for preference and perceived quality.
As an affiliate, we earn on qualifying purchases.
Key Questions
What does the Open TTS Leaderboard measure?
It reports word and character error rates for speech accuracy, offline generation speed, streaming time to first audio, and speaker similarity for voice cloning. These measures address different aspects of model performance and are not a single judgment of overall voice quality.
Does the leaderboard identify the most natural-sounding voice?
No. Hugging Face says automated metrics do not replace human preference rankings. The “Listen” tab lets users compare audio, but the announced material does not say that listener votes currently determine the default rankings.
How long does an evaluation take?
Hugging Face says evaluations using its objective metrics can take a couple of hours. It contrasts that with arena voting, which it says can take weeks; these are the project’s stated estimates.
Which models does Hugging Face name as leaders?
For multilingual performance, it names k2-fsa/OmniVoice, fishaudio/s2-pro and FunAudioLLM/Fun-CosyVoice3-0.5B-2512. For English error rates, it lists hexgrad/Kokoro-82M, Supertone/supertonic-3 and fishaudio/s2-pro. Those examples reflect selected leaderboard measures, not a universal quality verdict.
Will community votes change the rankings?
Hugging Face says votes may be included later, but it has not specified when or how they would affect rankings. Users can vote through the listening feature and are asked to sign in.
Primary source: Hugging Face · via ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
