Awesome Speaker Diarization
Section: Framework · SpeechBrain is an open-source and all-in-one speech toolkit based on PyTorch.
Entry
Appears in 5 awesome lists
PyTorch speech toolkit with recipes for ASR, TTS, speaker recognition, and speech enhancement.
Section: Framework · SpeechBrain is an open-source and all-in-one speech toolkit based on PyTorch.
Section: Speech · A PyTorch-based speech toolkit.
Section: 2. Model Codebases & Model Families · PyTorch speech toolkit with recipes for ASR, TTS, speaker recognition, and speech enhancement.
Section: Other · A PyTorch-based Speech Toolkit
Section: NLP & Speech Processing: · SpeechBrain is an open-source and all-in-one speech toolkit based on PyTorch.
Whisper is a general-purpose speech recognition model that can be run locally offline. It can transcribe audio from and to multiple languages.
industrial-grade ASR toolkit; 170× realtime on GPU, 50+ languages, built-in VAD, punctuation, speaker diarization, and emotion detection. Includes non-autoregressive SenseVoice and LLM-based Fun-ASR-Nano models.
VibeVoice is a novel framework designed for generating expressive, long-form, multi-speaker conversational audio, such as podcasts, from text. It addresses significant challenges in traditional Text-to-Speech (TTS) systems, particularly in scalability, speaker consistency, and natural turn-taking.
Kaldi is a toolkit for speech recognition written in C++ and licensed under the Apache License v2.0. Kaldi is intended for use by speech recognition researchers.
SDK for real-time on-device audio intelligence on iOS/macOS (diarization, identification, VAD, separation, embeddings, ASR), with CoreML models converted directly from PyTorch to leverage Apple Neural Engine performance.
OpenAI open-weight model repository with inference examples, recipes, and deployment guidance.
Speech-to-text, text-to-speech, speaker diarization, speech enhancement, source separation, and VAD using next-gen Kaldi with onnxruntime without Internet connection. Support embedded systems, Android, iOS, HarmonyOS, Raspberry Pi, RISC-V, RK NPU, Axera NPU, Ascend NPU, x86_64 servers, websocket…