Examples & Models

Kokoro, Qwen, F5, Higgs, and Chatterbox are the supported speech backends. Each model narrates the identical two-paragraph Gettysburg excerpt so you can directly compare audio timbre, pacing, pronunciation, and latency.

Speech Model Comparison

Compare five open model architectures running on Apple Silicon (MLX) and AWS Batch:

Kokoro-82M

by Hexgrad
82M params

The two-paragraph Gettysburg excerpt, spoken with Kokoro-82M on Apple Silicon (MLX) or PyTorch (~23s).

StyleTTS 2 + ISTFTNet24 kHzApache-2.0

Qwen3-TTS

by Alibaba Qwen
0.6B params

The same Gettysburg excerpt, spoken with Qwen3-TTS 0.6B CustomVoice with the Ryan preset (~48s).

CustomVoice 12Hz Transformer24 kHzApache-2.0

F5-TTS

by SWivid
Flow matching

The same Gettysburg excerpt, spoken with non-autoregressive F5-TTS via MLX on Apple Silicon.

Non-autoregressive Flow Matching24 kHzMIT

Higgs Audio v3

by Boson AI
Boson AI 4B

The same Gettysburg excerpt, spoken with Boson AI's expressive 4B Higgs Audio v3 via MLX on Apple Silicon.

Higgs Audio 4B Transformer24 kHzResearch / Non-Commercial

Chatterbox-TTS

by Resemble AI
520M Hybrid

The same Gettysburg excerpt, spoken with Resemble AI's hybrid LLaMA-520M and Matcha-TTS flow matching.

LLaMA-520M + Matcha-TTS Flow24 kHzApache-2.0

Fish-Speech

by Fish Audio
Dual-AR

The same Gettysburg excerpt, spoken with Fish Audio's dual-autoregressive model with natural prosody and breathing.

Dual-AR Transformer + VQ-GAN44.1 kHzCC-BY-NC-SA-4.0

Embed Features & Theming

Customizing player appearance and article content extraction: