Instructions to use Shoof/Gust-MLX with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use Shoof/Gust-MLX with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] hf download Shoof/Gust-MLX --local-dir Gust-MLX
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
Gust-MLX: real-time Breeze TTS 2 on Apple Silicon
Derived from Breeze TTS 2 by BreezeBlue and licensed for research and non-commercial use only.
An unofficial MLX conversion of BreezeBlue/Breeze-TTS-2, not affiliated with or endorsed by BreezeBlue, made for real-time streaming on Apple Silicon.
Modifications: the depth decoder's linear layers are quantized to int8 (affine, group size 64); everything else is unchanged from the bf16 MLX conversion.
The backbone, the text encoder and every embedding stay bf16, exactly as in mlx-community/Breeze-TTS-2-mlx. Only the depth decoder's linear layers are int8 (affine, group size 64). The depth decoder runs 15 times per audio frame and is about three quarters of the bytes each frame reads, so this is where quantization buys speed, while the backbone, which sets prosody and is where a voice direction lands, keeps its full precision.
Why
| Build (Apple M4 Pro, 48 GB) | First audio, plain / directed | RTF, directed | Memory |
|---|---|---|---|
BreezeBlue's PyTorch, bf16, on the Mac GPU (model.generate; its streaming path is CUDA-only) |
no streaming: a one-sentence line takes ~6 s | 2.26 (plain lines) | 11 GB |
| mlx-community 8-bit (mxfp8) | 212 / 297 ms | 0.58 | 13 GB |
| mlx-community bf16 | 380 / 456 ms | 0.86 | 15 GB |
| this checkpoint, with the streaming server's speed settings | 139 / 207 ms | 0.62 | 8.9 GB |
| For reference: RTX 3090 (power-capped at 250 W of 350 W), BreezeBlue's CUDA fast path via a small server wrapper, re-processing the reference on every line | 87 / 160 ms | 0.50 | 19.7 GB VRAM + 4 GB RAM |
On the same Mac, BreezeBlue's own PyTorch code runs at RTF 2.26 (over twice slower than real time) and can't stream; this build is ~3.6x faster and streams its first audio in 139 / 207 ms. For scale, a power-capped RTX 3090 running BreezeBlue's CUDA fast path measured 87 / 160 ms in our setup. The MLX builds' memory includes MLX's buffer cache (uncapped in the mlx-community builds, capped at 1 GB here).
Quality against bf16 on the same tokens (KL of the samplers' logits, lower is closer; lines with a voice direction at CFG 2):
| Build | Backbone KL | Depth decoder KL |
|---|---|---|
| mlx-community 8-bit (mxfp8) | 0.0030 | 0.0067 |
| this checkpoint | 0 (exact) | 0.0020 |
Voice directions amplify quantization error because CFG multiplies the gap between its two rows; the 8-bit build quantizes the backbone, this one doesn't. In a blind listening test (one listener, ten lines), the 8-bit build was judged worst on seven lines and best on none, while bf16 and this mixed build tied.
Measured with a 24 s LibriSpeech reference (speaker 1272, CC BY 4.0) and 24 lines of Harvard sentences, two in three with a voice direction. "First audio" includes any buffer playback needs so it never stalls.
Samples
The same text, voice direction (CFG 2) and seed on both. Voices cloned from LibriSpeech dev-clean readers 6295, 652, 2035, 84 (CC BY 4.0; Panayotov et al., 2015), about 20 s of reference each. Differences between the two are within normal sampling variation: in our listening, often neither is clearly better.
Male 1: "Speak warmly and gently, with a soft smile in the voice"
The birch canoe slid on the smooth planks. Glue the sheet to the dark blue background.
| RTX 3090 路 BreezeBlue's code | Gust-MLX 路 M4 Pro |
|---|---|
Male 2: "Speak in a restrained, serious, matter-of-fact tone"
It's easy to tell the depth of a well. These days a chicken leg is a rare dish.
| RTX 3090 路 BreezeBlue's code | Gust-MLX 路 M4 Pro |
|---|---|
Female 1: "Speak playfully, with a teasing, amused lilt"
Rice is often served in round bowls. The juice of lemons makes fine punch.
| RTX 3090 路 BreezeBlue's code | Gust-MLX 路 M4 Pro |
|---|---|
Female 2: "Whisper softly and intimately"
The box was thrown beside the parked truck. The hogs were fed chopped corn and garbage.
| RTX 3090 路 BreezeBlue's code | Gust-MLX 路 M4 Pro |
|---|---|
Use
With the streaming server from Shoofio/breeze-tts2-fast-streaming-api (HTTP and WebSocket APIs, voice cloning, voice direction):
uvx --from huggingface_hub hf download Shoof/Gust-MLX --local-dir Gust-MLX
BREEZE_MLX_COMPILE=1 BREEZE_MLX_FAST_FIRST=6 BREEZE_MLX_CACHE_GB=1 \
uv run python -m breeze_infer.api Gust-MLX --chunk-first 1 --chunk-max 1
(or scripts/start_breeze_mac.sh --precision mixed, which builds the same model from the bf16 weights at load).
It is a standard MLX checkpoint: mlx-audio loads it with its Breeze model as is, quantizing exactly the layers whose weights come with scales. If you stream with mlx-audio, note Blaizzy/mlx-audio#1003: mlx-audio versions without it add a bias twice at every chunk boundary in streaming codec decode (the server above patches it).
How it was made
scripts/make_mixed_checkpoint.py in the server repository: load the bf16 checkpoint with mlx-audio, quantize the depth decoder's linear layers (mlx.nn.quantize, 8 bits, affine, group 64), save, then load the result back with stock mlx-audio and compare every tensor with the in-memory conversion (identical). The tokenizer, the codec (audio_tokenizer/, the Qwen3-TTS 12 Hz tokenizer, Apache-2.0), LICENSE and NOTICE are copied unchanged.
License
Derived from Breeze-TTS-2 by BreezeBlue: the BreezeBlue Research and Non-Commercial License applies (research and non-commercial use only); see LICENSE and NOTICE. Read BreezeBlue's responsible-use terms before cloning anyone's voice.
Converted and measured with Claude Opus 5.5. Questions or results on other Macs welcome in the Community tab.
- Downloads last month
- 29
8-bit
Model tree for Shoof/Gust-MLX
Base model
BreezeBlue/Breeze-TTS-2