AI Game DevTools (AI-GDT)
Section: Speech · A multi-voice TTS system trained with an emphasis on quality.
Entry
Appears in 5 awesome lists
A multi-voice TTS system trained with an emphasis on quality github | research paper | demo
Section: Speech · A multi-voice TTS system trained with an emphasis on quality.
Section: Speech · A multi-voice text-to-speech system trained with an emphasis on quality. #opensource
Section: Text-to-speech (TTS) and avatars · "A multi-voice TTS system trained with an emphasis on quality"
Section: Repositories · A multi-voice TTS system trained with an emphasis on quality github | research paper | demo
Section: Text-to-speech · A multi-voice text-to-speech system trained with an emphasis on quality. #opensource
Whisper is a general-purpose speech recognition model that can be run locally offline. It can transcribe audio from and to multiple languages.
Bark is a transformer-based text-to-audio model created by Suno. Bark can generate highly realistic, multilingual speech as well as other audio - including music, background noise and simple sound effects.
VibeVoice is a novel framework designed for generating expressive, long-form, multi-speaker conversational audio, such as podcasts, from text. It addresses significant challenges in traditional Text-to-Speech (TTS) systems, particularly in scalability, speaker consistency, and natural turn-taking.
Kitten TTS is an open-source realistic text-to-speech model with just 15 million parameters, designed for lightweight deployment and high-quality voice synthesis.
🐸💬 - a deep learning toolkit for Text-to-Speech, battle-tested in research and production