2026 TTS Open-Source Boom: Open Models Approach Commercial Solutions
A May 2026 evaluation by Agora shows open-source models like CosyVoice 2 and Kokoro-82M achieving MOS scores close to Azure TTS in Chinese scenarios.
Read Original →Latest TTS developments, product launches, and industry highlights.
A May 2026 evaluation by Agora shows open-source models like CosyVoice 2 and Kokoro-82M achieving MOS scores close to Azure TTS in Chinese scenarios.
Read Original →Boson AI open-sources Higgs Audio V2, trained on 10 million hours of data, supporting emotion recognition, zero-shot voice cloning, and multi-character dialogue.
Read Original →VoxCPM2 adopts a Tokenizer-Free architecture, enabling multilingual speech generation and realistic voice cloning.
Read Original →Microsoft releases VibeVoice-1.5B under MIT license, supporting up to 90 minutes of continuous speech generation and 4-speaker concurrent output.
Read Original →Alibaba Cloud's Tongyi Lab Qwen team open-sources the Qwen3-TTS model series, supporting voice cloning, voice design, streaming synthesis, and natural language instruction control across 10 languages.
Read Original →Kokoro-82M delivers high-quality multilingual TTS with only 82M parameters under Apache 2.0, requiring just 4GB VRAM to run.
Read Original →HKUST and Mobvoi open-source Spark-TTS, a zero-shot voice cloning model built on Qwen2.5 with 0.5B parameters for high-quality speech synthesis.
Read Original →Alibaba Tongyi Lab's CosyVoice 2 achieves an MOS of 4.7 in Chinese TTS evaluation, supports Cantonese and Shanghainese, and is open-sourced under Apache 2.0.
Read Original →CosyVoice 2.0 achieves significant breakthroughs in streaming speech synthesis, zero-shot voice cloning, and emotion control, now open-sourced on GitHub.
Read Original →Fish Speech 1.4 delivers significant improvements in multi-speaker timbre consistency, Chinese prosody naturalness, and inference speed.
Read Original →GPT-4o's real-time voice mode supports end-to-end voice conversations with latency reduced to 200ms, marking a new phase in AI voice interaction.
Read Original →ChatTTS continues to attract massive attention on GitHub, with its conversational speech synthesis leading the development of Chinese open-source TTS.
Read Original →OpenAI integrates native TTS capabilities into GPT-4o, supporting multilingual speech synthesis and emotional expression, with the API now generally available.
Read Original →