Skip to main content

Spark-TTS

GitHub ⭐ 11000License: Apache 2.0Model Type: Zero-shot TTS / Voice Cloning

A lightweight zero-shot voice cloning model jointly developed by HKUST and Mobvoi, built on Qwen2.5 with only 0.5B parameters for high-quality speech synthesis.

ChineseEnglish
Zero-shotVoice CloningHKUSTRecommended

Spark-TTS is an efficient text-to-speech system jointly developed by the Hong Kong University of Science and Technology (HKUST) and Mobvoi (出门问问), in collaboration with Shanghai Jiao Tong University, Nanyang Technological University, and Northwestern Polytechnical University. Led by the SparkAudio team, the project’s paper was published on arXiv (2503.01710). Spark-TTS is built entirely on the Qwen2.5 large language model with only 0.5B parameters, eliminating the need for additional generation modules such as flow matching that are common in traditional TTS systems. By using Single-Stream Decoupled Speech Tokens, it directly reconstructs audio from the codes predicted by the LLM, achieving a remarkably streamlined and efficient architecture.

In terms of capabilities, Spark-TTS supports zero-shot voice cloning — requiring only a few seconds of reference audio to accurately replicate a target speaker’s timbre, and can handle cross-lingual and code-switching scenarios. The model supports bilingual synthesis in both Chinese and English. Additionally, Spark-TTS offers controllable speech generation, allowing users to create virtual speakers by adjusting parameters such as gender, pitch, and speaking rate without any reference audio. This “Voice Creation” capability holds unique practical value for content creation and personalized voice assistant applications.

Spark-TTS is released under the Apache 2.0 open-source license and has garnered over 11,000 stars on GitHub. Model weights are available on Hugging Face (SparkAudio/Spark-TTS-0.5B), and the project provides complete inference code, a Gradio WebUI, and Nvidia Triton inference serving deployment solutions. With its lightweight design, zero-shot cloning capability, and permissive open-source license, Spark-TTS demonstrates broad application potential in academic research, personalized speech synthesis, assistive technologies, and cross-lingual voice applications.

Tags:Zero-shotVoice CloningHKUSTRecommended
Share:X