AI Game DevTools (AI-GDT)
Section: Speech Β· XTTS is a library for advanced Text-to-Speech generation.
Entry
Appears in 4 awesome lists
πΈπ¬ - a deep learning toolkit for Text-to-Speech, battle-tested in research and production
Section: Speech Β· XTTS is a library for advanced Text-to-Speech generation.
Section: Speech and Text Β· and VieNeu-TTS - open TTS.
Section: Natural Language Processing Β· πΈπ¬ - a deep learning toolkit for Text-to-Speech, battle-tested in research and production
Section: Acoustic User Interface Β· A deep learning toolkit for Text-to-Speech, battle-tested in research and production.
Whisper is a general-purpose speech recognition model that can be run locally offline. It can transcribe audio from and to multiple languages.
industrial-grade ASR toolkit; 170Γ realtime on GPU, 50+ languages, built-in VAD, punctuation, speaker diarization, and emotion detection. Includes non-autoregressive SenseVoice and LLM-based Fun-ASR-Nano models.
Industrial-strength natural language processing with 75+ languages, transformer pipelines, and production-grade NER, parsing, and text classification.
Bark is a transformer-based text-to-audio model created by Suno. Bark can generate highly realistic, multilingual speech as well as other audio - including music, background noise and simple sound effects.
Stanford NLP Python library for 100+ human languages. State-of-the-art neural pipelines for tokenization, NER, parsing, and sentiment analysis with pre-trained models. Apache 2.0 licensed.
VibeVoice is a novel framework designed for generating expressive, long-form, multi-speaker conversational audio, such as podcasts, from text. It addresses significant challenges in traditional Text-to-Speech (TTS) systems, particularly in scalability, speaker consistency, and natural turn-taking.