Skip to main content

Higgs Audio V2 Open-Sourced: Mu Li's Team Brings the Most Expressive Emotional TTS

Boson AI open-sources Higgs Audio V2, trained on 10 million hours of data, supporting emotion recognition, zero-shot voice cloning, and multi-character dialogue.

·Boson AI
Read Original →

Higgs Audio V2 is currently the most emotionally expressive open-source TTS model, developed by Mu Li’s team. It uses a unified audio tokenizer with residual vector quantization, achieving 2kbps compression while maintaining 24kHz quality. Chinese MOS 4.6, English MOS 4.5, and emotional expression MOS 4.7 — the highest in the industry.

Zero-shot cloning requires only 5 seconds of reference audio. Recommended use cases: virtual streamers, conversational assistants, audiobooks.

The model received a 5-star recommendation in Agora’s May 2026 evaluation and stands out as the most prominent open-source TTS model for emotional expression.

Share:X