|

Breeze TTS 2 brings open-weight voice generation closer to the best proprietary models

Breeze TTS 2 Leaderboard

BreezeBlue has released Breeze TTS 2, a 3-billion-parameter text-to-speech model designed for real-time voice applications. Available as downloadable weights, it combines voice design, voice cloning, performance direction and streaming generation in a single system.

The model is particularly relevant for voice agents, game characters, interactive storytelling and other applications that require more than conventional text-to-speech. A voice can be created from a natural-language description, then directed through instructions controlling its emotion, tone, pace and delivery. Breeze TTS 2 can also reproduce a voice from a reference recording and its exact transcript, while supporting vocal events such as laughter, sighs and coughs.

In the video below, we present the model, its main capabilities, its position among current TTS systems and the important limitations of its local license.

Breeze TTS 2 ranks sixth in the Artificial Analysis Speech Arena

Continue reading after the ad

At the time of writing, Breeze TTS 2 holds sixth place in the Artificial Analysis Provider Voice leaderboard, with an Elo score of approximately 1,215. It is also the highest-ranked open-weight model in the benchmark, ahead of Fish Audio S2 Pro, Step Audio EditX, Voxtral TTS and Kokoro.

The ranking is based on blind listening comparisons rather than benchmarks published solely by the model developer. It does not make Breeze TTS 2 the best option for every project, but it places the model unusually close to leading proprietary services. This makes it an interesting option for developers who want to experiment locally with expressive voices, character creation and low-latency speech generation.

BreezeBlue reports a time to first audio below 40 milliseconds and a real-time factor of 0.32 on an Nvidia H100 using its optimized fast path. These figures describe a specific high-end configuration and should not be treated as expected performance on every local GPU.

The local model is restricted to non-commercial use

Breeze TTS 2 should be described as an open-weight model rather than fully open source. Its PyTorch inference code is published under the Apache 2.0 license, but the downloaded model weights are governed by the BreezeBlue Research and Non-Commercial License.

This distinction also applies to generated content. According to the Breeze TTS 2 model page, the weights, derivative models and outputs produced through a self-hosted instance are restricted to research and non-commercial use. Commercial use requires written authorization from RESONIA, INC. or access through a commercial BreezeBlue service under the applicable terms.

Continue reading after the ad

As a result, the local checkpoint should not be used by default for monetized videos, paid client work, commercial games or revenue-generating applications. Downloadable weights do not automatically provide commercial rights.

Early benchmarks: speaker similarity vs. text stability

To evaluate how Breeze TTS 2 performs outside standard vendor leaderboards, early comparative testing provides valuable insight into its text accuracy and voice cloning fidelity compared to other leading open-weight architectures.

The benchmark below evaluates Word Error Rate (WER) and Character Error Rate (CER), where lower values indicate higher text stability, alongside Speaker Similarity (SIM), where higher percentages reflect closer acoustic matching.

ModelParametersEnglish (WER / SIM)Chinese (CER / SIM)
Audio8 0.1B0.17B1.662 / 56.7%1.130 / 68.2%
Audio8 0.6B0.6B1.506 / 63.2%0.950 / 73.1%
OmniVoice0.61B1.620 / 74.0%0.870 / 77.7%
Confucius4 TTS0.77B1.490 / 70.0%0.940 / 76.5%
IndexTTS 2.50.8B3.253 / 82.3%1.119 / 80.4%
CosyVoice 31.5B2.220 / 72.0%1.120 / 78.1%
Qwen3 TTS1.7B1.240 / 71.4%0.770 / 77.0%
VoxCPM 22.3B1.700 / 75.2%0.970 / 79.3%
Breeze TTS 23.5B2.143 / 68.2%1.465 / 74.5%
Fish Audio S24.6B1.790 / 64.3%0.980 / 73.7%
Higgs Audio v24.7B1.524 / 66.4%0.806 / 72.1%
MOSS TTS8.5B1.850 / 73.4%1.200 / 78.8%
Source: Official Discord “Breeze Blue Studio

Analyzing the trade-offs: raw metrics vs. expressive direction

At 3.5B parameters, Breeze TTS 2 holds a respectable speaker similarity score, peaking at 74.5% in Chinese and 68.2% in English, performing competitively against mid-sized counterparts. However, its text stability scores (WER of 2.143 and CER of 1.465) highlight that strict verbatim consistency and hallucination-free generation remain areas with room for refinement compared to ultra-compact models like Qwen3 TTS or Higgs Audio v2.

Continue reading after the ad

Yet, raw WER and SIM metrics tell only part of the story. Standard objective benchmarks measure strict acoustic cloning and transcription fidelity, but they completely overlook:

  • Zero-shot voice design capabilities.
  • Natural language voice direction and prompt adherence.
  • Expressive delivery and nuanced paralinguistic cues (laughter, sighs, breath control).
  • Overall human preference and perceived organic prosody.

While models optimized purely for low WER deliver reliable transcription accuracy, Breeze TTS 2 shifts the paradigm toward director-level control and emotive expressiveness, making it one of the most innovative and versatile open-weight voice models released to date, but without a commercial license.

A strong model, but not yet the fully open TTS alternative

There is currently no clearly permissive open-source TTS model for commercial local use that matches Breeze TTS 2 across this entire combination of voice design, voice cloning, natural-language direction, expressive events, streaming and competitive listening quality.

That gap remains surprising. Advanced open-weight generation models such as MiniMax H3 show that complex multimodal systems can run locally and provide a pathway to commercial use under defined conditions. H3 does, however, impose licensing conditions and requires formal authorization for local deployment in regions including the European Union.

TTS development may appear more accessible than training a large multimodal video model, yet local speech synthesis still lacks a permissively licensed system that clearly rivals ElevenLabs across quality, control and features—or consistently comes close to the complete service.

Breeze TTS 2 therefore represents meaningful progress and a compelling model for research, testing and non-commercial local projects. Its technical capabilities are open for developers to explore, but its license prevents it from becoming the unrestricted commercial alternative many creators are still waiting for.

Your comments enrich our articles, so don’t hesitate to share your thoughts! Sharing on social media helps us a lot. Thank you for your support!

Continue reading after the ad

Similar Posts

Leave a Reply

Your email address will not be published. Required fields are marked *