Open TTS Leaderboard: Objective Evaluation for Multilingual TTS Models

Overview of the Open TTS Leaderboard

Launched on September 30, 2026, the Open TTS Leaderboard introduces a scalable and objective approach to evaluating text-to-speech (TTS) models, addressing the limitations of traditional vote-based arena rankings. With over 8,000 TTS models available on the Hugging Face Hub, the need for a standardized evaluation method has become increasingly apparent.

Objective Metrics and Methodology

The leaderboard employs automated metrics to assess models across several performance dimensions. Intelligibility is gauged using word error rate (WER) or character error rate (CER), comparing the original prompt with an ASR transcript generated by the Qwen3 ASR model. Speed is evaluated through the inverse real-time factor (RTFx) for batched offline inference on an H200 GPU, as well as time-to-first-audio (TTFA) for streaming scenarios on both H200 GPU and CPU. Additionally, speaker similarity (SIM) is calculated by measuring the cosine similarity between WavLM speaker embeddings of the generated audio and a reference clip.

It is important to note that these automated metrics serve as proxies for more subjective measures. While ASR-based WER/CER provides insights into intelligibility, SIM offers an estimate of identity preservation. However, they do not replace human preference evaluations, such as Mean Opinion Score (MOS) or MUSHRA, which remain critical in determining the overall quality of TTS outputs.

Multilingual Performance and Voice Cloning

By default, models are ranked based on macro-average WER from English splits of Seed TTS Eval and CV3 Eval (zero-shot). Leading models in this category include hexgrad/Kokoro-82M, Supertone/supertonic-3, and fishaudio/s2-pro. The leaderboard also features Pareto plots to illustrate the trade-offs between WER, batched inference (RTFx), and model size.

Users can toggle multilingual rankings, with Seed TTS Eval providing audio for English and Chinese, while other languages are assessed using CV3 Eval. For character-based languages such as Chinese, Japanese, and Korean, CER is reported, and an average WER is calculated as a macro-average across all languages. Notable multilingual models include k2-fsa/OmniVoice, fishaudio/s2-pro, and FunAudioLLM/Fun-CosyVoice3-0.5B-2512. The leaderboard also includes a voice cloning feature, which adds a SIM column and additional Pareto visualizations. Some models, like bosonai/higgs-tts-3-4b and openbmb/VoxCPM2, demonstrate improved average WER when a reference audio is provided.

Listening and Voting Features

The “Listen” tab allows users to compare generated outputs from different models and cast votes based on their preferences. This feature fills a significant gap in existing TTS leaderboards by providing a platform for exploring model outputs. Users are encouraged to log in with their Hugging Face accounts to help mitigate spam and ensure the integrity of the voting process. As community votes are collected, there is potential for this data to be integrated into the leaderboard over time.

Streaming Performance Evaluation

The “Streaming” tab ranks models based on TTFA, measuring the time it takes for the first playable audio chunk to be generated for streaming models, or the total time for non-streaming models to complete utterance generation. All tests are conducted using identical prompts, with initial warm-up runs excluded from the results. The default evaluation is performed on an H200 GPU, with a growing number of CPU results available. The model kyutai/pocket-tts is highlighted for its low latency performance on both GPU and CPU.

Future Developments and Community Involvement

As of now, only 16 out of 92 models on the Artificial Analysis leaderboard are open-weight, highlighting the challenges faced by open-source models in gaining visibility. The Open TTS Leaderboard aims to address this by focusing on open-source models and multilingual capabilities. The authors plan to open-source the evaluation scripts in the near future, inviting community feedback through GitHub Issues and pull requests to enhance the evaluation process.


Original source: Hugging Face Blog

Leave a Comment