VoxCPM2
OpenBMB's VoxCPM2 tokenizer-free speech synthesis model — 2B parameters, 30 languages and 8 dialects, 48kHz studio-grade output, with voice editing and voice design capabilities.
VoxCPM2 is the second-generation tokenizer-free speech synthesis model developed by OpenBMB, innovatively modeling raw audio waveforms directly and eliminating the intermediate tokenizer-dependent steps found in traditional TTS systems. With 2B parameters, this end-to-end design avoids the accumulation of tokenization errors, allowing the model to capture subtle features in speech signals more accurately and produce exceptionally natural synthesized speech.
In terms of capability coverage, VoxCPM2 supports up to 30 languages and 8 major dialects, making it one of the broadest-coverage open-source TTS models available. The model outputs studio-quality audio at 48kHz, meeting the demands of professional audio production. Beyond voice cloning and controllable emotional expression, VoxCPM2 introduces voice editing and voice design capabilities — users can finely edit existing speech or design entirely new voice characteristics from scratch, greatly expanding creative possibilities.
VoxCPM2 is open-sourced under the Apache 2.0 license, making it commercially friendly. The project has received over 12,000 stars on GitHub with a highly active and growing community. Whether you need cross-lingual speech synthesis, personalized voice customization, or emotionally expressive speech generation and editing, VoxCPM2 provides a powerful technical foundation. For more information, visit the VoxCPM GitHub Repository.