Skip to main content

VibeVoice

GitHub ⭐ 6200License: MITModel Type: Long-Form Audio Generation TTS

Microsoft's open-source project, originally a 1.5B long-form TTS model. TTS training/inference code removed as of Sep 2025; project now focused on ASR and speech processing.

ChineseEnglish
MicrosoftLong-Form AudioMITMulti-SpeakerASR

VibeVoice is a 1.5B parameter large-scale TTS model open-sourced by Microsoft Research, purpose-built for long-form audio generation. Unlike traditional TTS models, VibeVoice can generate up to 90 minutes of continuous speech in a single inference pass, eliminating the need for segment stitching and thus preserving naturalness and coherence over long text.

Architecturally, VibeVoice employs continuous speech tokenization combined with a diffusion generation architecture, overcoming the bottlenecks that autoregressive models face in long-sequence generation. The model supports timbre switching for up to 4 speakers, making it suitable for content creation scenarios that require multi-role delivery. Its long-form audio consistency MOS reaches 4.6, significantly outperforming comparable models.

⚠️ Project Status Update: In September 2025, Microsoft removed the TTS training and inference code from the VibeVoice repository. The project has since pivoted to focus on ASR (Automatic Speech Recognition) and speech processing. The original TTS model weights and paper remain available, but are no longer maintained through this repository. Check the VibeVoice GitHub Repository for the latest status.

Tags:MicrosoftLong-Form AudioMITMulti-SpeakerASR
Share:X