Skip to content

Entry

sherpa-onnx

Appears in 4 awesome lists

Speech-to-text, text-to-speech, speaker diarization, speech enhancement, source separation, and VAD using next-gen Kaldi with onnxruntime without Internet connection. Support embedded systems, Android, iOS, HarmonyOS, Raspberry Pi, RISC-V, RK NPU, Axera NPU, Ascend NPU, x86_64 servers, websocket…

Open github.comk2-fsa/sherpa-onnx

Found in these lists

Awesome Speaker Diarization

Section: Framework · Support speaker diarization, speech recognition, and text-to speech on various platforms with various language bindings.

FreshScore 84

Awesome Open Source AI

Section: 2. Model Codebases & Model Families · Complete speech toolkit with ASR, TTS, diarization, source separation, and VAD across embedded and edge environments via ONNX Runtime.

FreshScore 89

Awesome Pascal

Section: Machine Learning · . [Delphi] [FPC] Speech-to-text, text-to-speech, speaker diarization, speech enhancement, source separation, and VAD using next-gen Kaldi with onnxruntime without Internet connection. Support embedded systems, Android, iOS, HarmonyOS, Raspberry Pi, RISC-V, RK NPU, Ascend NPU, x86_64 servers,…

SlowScore 66

awesome-cpp

Section: Natural Language Processing · Speech-to-text, text-to-speech, speaker diarization, speech enhancement, source separation, and VAD using next-gen Kaldi with onnxruntime without Internet connection. Support embedded systems, Android, iOS, HarmonyOS, Raspberry Pi, RISC-V, RK NPU, Axera NPU, Ascend NPU, x86_64 servers, websocket…

FreshScore 79

Whisper

Whisper is a general-purpose speech recognition model that can be run locally offline. It can transcribe audio from and to multiple languages.

In 11 listsDetails

FunASR

industrial-grade ASR toolkit; 170× realtime on GPU, 50+ languages, built-in VAD, punctuation, speaker diarization, and emotion detection. Includes non-autoregressive SenseVoice and LLM-based Fun-ASR-Nano models.

In 9 listsDetails

VibeVoice

VibeVoice is a novel framework designed for generating expressive, long-form, multi-speaker conversational audio, such as podcasts, from text. It addresses significant challenges in traditional Text-to-Speech (TTS) systems, particularly in scalability, speaker consistency, and natural turn-taking.

In 5 listsDetails

Kaldi

Kaldi is a toolkit for speech recognition written in C++ and licensed under the Apache License v2.0. Kaldi is intended for use by speech recognition researchers.

In 5 listsDetails

SpeechBrain

PyTorch speech toolkit with recipes for ASR, TTS, speaker recognition, and speech enhancement.

In 5 listsDetails

FluidAudio

SDK for real-time on-device audio intelligence on iOS/macOS (diarization, identification, VAD, separation, embeddings, ASR), with CoreML models converted directly from PyTorch to leverage Apple Neural Engine performance.

In 4 listsDetails

gpt-oss

OpenAI open-weight model repository with inference examples, recipes, and deployment guidance.

In 4 listsDetails

MiniCPM-V

Compact vision-language model family with edge-focused deployment examples and strong OCR-oriented use cases.

In 4 listsDetails