TTS Open-Source Projects
A curated collection of TTS open-source projects on GitHub, updated regularly.
Bark
Suno AI's open-source Transformer-based text-to-audio model that can synthesize not only speech but also laughter, sighs, music, and background sound effects.
ChatTTS
A TTS model designed for conversational scenarios, supporting fine-grained prosody control and multi-speaker capabilities, continuously updated as of April 2026.
Coqui TTS
A well-known end-to-end TTS framework offering XTTS multilingual voice cloning, model training and fine-tuning, with a rich community ecosystem.
CosyVoice
Alibaba Tongyi Lab's open-source streaming TTS large model. CosyVoice 2 achieved a Chinese MOS of 4.7, topping benchmarks, with support for dialects and instruction control.
F5-TTS
A next-generation non-autoregressive TTS model based on Flow Matching, offering extremely fast inference and excellent audio quality.
Fish Speech
A TTS framework based on VQ-GAN and language models. The latest S2 Pro version features 4B parameters, supports 80+ languages, and is trained on 10 million hours of data.
GPT-SoVITS
A milestone project in few-shot voice cloning, combining GPT and SoVITS architectures, capable of cloning a voice with just 1 minute of audio.
Kokoro-82M
A high-performance multilingual TTS model with only 82M parameters, Apache 2.0 licensed for commercial use, running on just 4GB VRAM.
MeloTTS
MyShell's open-source lightweight Chinese TTS model, supporting real-time inference on CPU, ideal for edge deployment.
OpenVoice
MyShell's open-source voice cloning framework supporting fine-grained timbre and style control, capable of cloning a voice from just a short audio sample.
Qwen3-TTS
An open-source multilingual speech synthesis model series from Alibaba Cloud's Qwen team, supporting voice cloning, voice design, streaming generation, and natural language instruction control.
Spark-TTS
A lightweight zero-shot voice cloning model jointly developed by HKUST and Mobvoi, built on Qwen2.5 with only 0.5B parameters for high-quality speech synthesis.
VibeVoice
Microsoft's open-source project, originally a 1.5B long-form TTS model. TTS training/inference code removed as of Sep 2025; project now focused on ASR and speech processing.
VITS
A classic end-to-end TTS model with a single-stage VAE architecture that directly converts text to waveform, serving as the foundation for many subsequent TTS projects.
VoxCPM2
OpenBMB's VoxCPM2 tokenizer-free speech synthesis model — 2B parameters, 30 languages and 8 dialects, 48kHz studio-grade output, with voice editing and voice design capabilities.