AI Game DevTools (AI-GDT)
Section: Speech · VoxCPM: Tokenizer-Free TTS for Context-Aware Speech Generation and True-to-Life Voice Cloning.
Entry
Appears in 4 awesome lists
Open-sourced tokenizer-free multilingual speech synthesis model with high-quality TTS and style transfer workflows.
Section: Speech · VoxCPM: Tokenizer-Free TTS for Context-Aware Speech Generation and True-to-Life Voice Cloning.
Section: 语音 Speech
Section: 2. Model Codebases & Model Families · Open-sourced tokenizer-free multilingual speech synthesis model with high-quality TTS and style transfer workflows.
Section: Natural Language Processing · VoxCPM2: Tokenizer-Free TTS for Multilingual Speech Generation, Creative Voice Design, and True-to-Life Cloning
Whisper is a general-purpose speech recognition model that can be run locally offline. It can transcribe audio from and to multiple languages.
industrial-grade ASR toolkit; 170× realtime on GPU, 50+ languages, built-in VAD, punctuation, speaker diarization, and emotion detection. Includes non-autoregressive SenseVoice and LLM-based Fun-ASR-Nano models.
Bark is a transformer-based text-to-audio model created by Suno. Bark can generate highly realistic, multilingual speech as well as other audio - including music, background noise and simple sound effects.
VibeVoice is a novel framework designed for generating expressive, long-form, multi-speaker conversational audio, such as podcasts, from text. It addresses significant challenges in traditional Text-to-Speech (TTS) systems, particularly in scalability, speaker consistency, and natural turn-taking.
A multi-voice TTS system trained with an emphasis on quality github | research paper | demo
PyTorch speech toolkit with recipes for ASR, TTS, speaker recognition, and speech enhancement.