Skip to main content

Industry News

Latest TTS developments, product launches, and industry highlights.

May 26, 2026·Agora

2026 TTS Open-Source Boom: Open Models Approach Commercial Solutions

A May 2026 evaluation by Agora shows open-source models like CosyVoice 2 and Kokoro-82M achieving MOS scores close to Azure TTS in Chinese scenarios.

Read Original →
March 15, 2026·Boson AI

Higgs Audio V2 Open-Sourced: Mu Li's Team Brings the Most Expressive Emotional TTS

Boson AI open-sources Higgs Audio V2, trained on 10 million hours of data, supporting emotion recognition, zero-shot voice cloning, and multi-character dialogue.

Read Original →
March 1, 2026·OpenBMB

OpenBMB Releases VoxCPM2: Tokenizer-Free Multilingual Speech Synthesis

VoxCPM2 adopts a Tokenizer-Free architecture, enabling multilingual speech generation and realistic voice cloning.

Read Original →
January 20, 2026·Microsoft Research

Microsoft Open-Sources VibeVoice-1.5B: Generates Up to 90 Minutes of Long-Form Audio in a Single Pass

Microsoft releases VibeVoice-1.5B under MIT license, supporting up to 90 minutes of continuous speech generation and 4-speaker concurrent output.

Read Original →
December 1, 2025·Alibaba Tongyi Lab

Alibaba Tongyi Open-Sources Qwen3-TTS: A New Benchmark in Multilingual Speech Synthesis

Alibaba Cloud's Tongyi Lab Qwen team open-sources the Qwen3-TTS model series, supporting voice cloning, voice design, streaming synthesis, and natural language instruction control across 10 languages.

Read Original →
November 20, 2025·hexgrad

Kokoro-82M: An Ultra-Lightweight Open-Source TTS Engine with Just 82M Parameters

Kokoro-82M delivers high-quality multilingual TTS with only 82M parameters under Apache 2.0, requiring just 4GB VRAM to run.

Read Original →
October 1, 2025·SparkAudio

Spark-TTS Open-Sourced: A Lightweight Voice Cloning Model from HKUST and Mobvoi

HKUST and Mobvoi open-source Spark-TTS, a zero-shot voice cloning model built on Qwen2.5 with 0.5B parameters for high-quality speech synthesis.

Read Original →
August 15, 2025·Alibaba Tongyi Lab

CosyVoice 2 Tops Chinese TTS Benchmarks: Alibaba Tongyi Lab's Open-Source Model Continues to Evolve

Alibaba Tongyi Lab's CosyVoice 2 achieves an MOS of 4.7 in Chinese TTS evaluation, supports Cantonese and Shanghainese, and is open-sourced under Apache 2.0.

Read Original →
September 1, 2024·Alibaba

Alibaba Tongyi Lab Open-Sources CosyVoice 2.0 with Major Advances in Streaming Voice Cloning

CosyVoice 2.0 achieves significant breakthroughs in streaming speech synthesis, zero-shot voice cloning, and emotion control, now open-sourced on GitHub.

Read Original →
August 1, 2024·Fish Audio

Fish Speech 1.4 Released: Chinese TTS Quality Reaches New Heights

Fish Speech 1.4 delivers significant improvements in multi-speaker timbre consistency, Chinese prosody naturalness, and inference speed.

Read Original →
July 1, 2024·OpenAI

OpenAI Launches GPT-4o Real-Time Voice Mode, Ushering in a New Era of TTS Interaction

GPT-4o's real-time voice mode supports end-to-end voice conversations with latency reduced to 200ms, marking a new phase in AI voice interaction.

Read Original →
June 15, 2024·GitHub

ChatTTS Continues to Dominate, Chinese TTS Open-Source Project Stays Hot

ChatTTS continues to attract massive attention on GitHub, with its conversational speech synthesis leading the development of Chinese open-source TTS.

Read Original →
May 13, 2024·OpenAI

OpenAI Launches GPT-4o TTS Capabilities

OpenAI integrates native TTS capabilities into GPT-4o, supporting multilingual speech synthesis and emotional expression, with the API now generally available.

Read Original →