Awesome Speaker Diarization
Section: Framework · A native Swift speaker diarization library for Apple platforms, using CoreML for efficient, real-time audio processing with high accuracy.
Entry
Appears in 4 awesome lists
SDK for real-time on-device audio intelligence on iOS/macOS (diarization, identification, VAD, separation, embeddings, ASR), with CoreML models converted directly from PyTorch to leverage Apple Neural Engine performance.
Section: Framework · A native Swift speaker diarization library for Apple platforms, using CoreML for efficient, real-time audio processing with high accuracy.
Section: Audio · SDK for real-time on-device audio intelligence on iOS/macOS (diarization, identification, VAD, separation, embeddings, ASR), with CoreML models converted directly from PyTorch to leverage Apple Neural Engine performance.
Section: Audio · Swift framework for local speech recognition, speaker diarization, voice activity detection, and text-to-speech using Core ML.
Section: Audio · On-device speech processing for iOS and macOS: ASR, TTS, VAD, and speaker diarization
industrial-grade ASR toolkit; 170× realtime on GPU, 50+ languages, built-in VAD, punctuation, speaker diarization, and emotion detection. Includes non-autoregressive SenseVoice and LLM-based Fun-ASR-Nano models.
Kaldi is a toolkit for speech recognition written in C++ and licensed under the Apache License v2.0. Kaldi is intended for use by speech recognition researchers.
PyTorch speech toolkit with recipes for ASR, TTS, speaker recognition, and speech enhancement.
Speech-to-text, text-to-speech, speaker diarization, speech enhancement, source separation, and VAD using next-gen Kaldi with onnxruntime without Internet connection. Support embedded systems, Android, iOS, HarmonyOS, Raspberry Pi, RISC-V, RK NPU, Axera NPU, Ascend NPU, x86_64 servers, websocket…
Python library for audio segmentation, silence removal via dynamic thresholding, spectral features, and classification.
Persistence player to resume playback after bad network connection even in background mode, manage headphone interactions, system interruptions, now playing informations and remote commands.
Powerful audio synthesis, processing and analysis, without the steep learning curve.
Neural building blocks for speaker diarization: speech activity detection, speaker change detection, speaker embedding.