BreezeBlue Breeze TTS 2 Open Weights
As of 14 Sep 2026, BreezeBlue Breeze TTS 2 Open Weights is the RightSignal pick for best open-weight TTS.
Top open-weight model on the Artificial Analysis provider-voice leaderboard at 1202 Elo, ninth overall and 82 points clear of the next open-weight model, Fish Audio S2 Pro at 1120.
Leads the pick 1207 to 1202 on the Artificial Analysis provider-voice board, seventh overall against ninth. It is proprietary and API-only on StepFun's platform at about $85 per 1M characters, though, with no weights published, so it doesn't displace the pick for self-hosters. StepFun's open-weight audio line remains Step-Audio-EditX and Step-Audio-2-mini.
Why it’s the challengercurrent reign
held off
reigns
Why BreezeBlue Breeze TTS 2 Open Weights is the pick
- Top open-weight model on the Artificial Analysis provider-voice leaderboard at 1202 Elo, ninth overall and 82 points clear of the next open-weight model, Fish Audio S2 Pro at 1120.
- BreezeBlue quotes under 40ms time-to-first-audio and a 0.32 RTF on a warmed H100, but those numbers come from the --fast-all path, which wants about 14.4 GiB and a 24GB card; plain eager inference fits in roughly 7.7 GiB, so 12GB is the practical floor.
- English and Chinese only, and the weights sit under the BreezeBlue Research and Non-Commercial Licence, so commercial self-hosting still needs a separate arrangement.
Evidence
The sources behind this title’s record.
- huggingface.co/BreezeBlue/Breeze-TTS-2
- huggingface.co/tencent/AuK
- artificialanalysis.ai/text-to-speech/arena?tab=leaderboard
- huggingface.co/fishaudio/s2-pro
- artificialanalysis.ai/text-to-speech
- artificialanalysis.ai/text-to-speech/leaderboard
- arxiv.org/abs/2608.26146
- arxiv.org/abs/2603.09215
Vendor numbers are treated as claims until independently reproduced — how we judge. Structured benchmark comparisons are on the roadmap.
Caveats & challengers
- Open-weight but not open-source: the inference code is Apache-2.0, but the weights, derivatives and self-hosted outputs fall under the BreezeBlue Research and Non-Commercial Licence, and a paid breezeblue.ai subscription only covers outputs from the hosted API — not anything you generate locally.
- That keeps Apache-2.0 Step Audio EditX (1099 Elo on the same board) the pragmatic fallback if you need commercial self-hosting.
- The lead is also board-specific and language-limited: on Artificial Analysis's controlled-voices arena Breeze ranks third among open-weight models at 1002, behind Mistral's Voxtral TTS at 1010, and the model speaks English and Chinese only.
Leads the pick 1207 to 1202 on the Artificial Analysis provider-voice board, seventh overall against ninth. It is proprietary and API-only on StepFun's platform at about $85 per 1M characters, though, with no weights published, so it doesn't displace the pick for self-hosters. StepFun's open-weight audio line remains Step-Audio-EditX and Step-Audio-2-mini.
At a glance
- Title
- Best open-weight TTS
- Licence
- Apache-2.0 inference code; weights under BreezeBlue research/non-commercial licence
- Pick since
- 9 Sep 2026
- Last reviewed
- 13 Sep 2026
- Title holders to date
- 8
- Official page
- huggingface.co/BreezeBlue/Breeze-TTS-2
Title history
Every change, on the record.
Top open-weight model on the Artificial Analysis provider-voice leaderboard at 1202 Elo, ninth overall and 82 points clear of the next open-weight model, Fish Audio S2 Pro at 1120.
5 daysas pick
80+ language coverage, fine-grained control and genuinely usable streaming latency; 550k+ downloads say it displaced everything else as t…
184 daysas pick
Apache-2.0 weights that added iterative post-hoc editing of emotion and style on top of strong zero-shot cloning — the workflow feature p…
117 daysas pick
The first open model with credible blind-test evidence of beating ElevenLabs (63.75% preference), shipped MIT with emotion-exaggeration c…
168 daysas pick
An 82M-parameter model that sounded better than things twenty times its size and ran on CPU — roughly 12M downloads and the default for a…
153 daysas pick
Flow matching gave noticeably cleaner prosody and faster inference than XTTS; the community moved on quickly despite the non-commercial l…
80 daysas pick
Made six-second zero-shot voice cloning across 17 languages actually work, and became the workhorse behind most open TTS products for the…
342 daysas pick
as pick