AI Game DevTools (AI-GDT)
Section: Speech · Whisper is a general-purpose speech recognition model.
Entry
Appears in 11 awesome lists
Whisper is a general-purpose speech recognition model that can be run locally offline. It can transcribe audio from and to multiple languages.
Section: Speech · Whisper is a general-purpose speech recognition model.
Section: Others · Robust Speech Recognition via Large-Scale Weak Supervision
Section: Tools · Robust speech recognition model for transcription and translation.
Section: 语音 Speech
Section: Audio Foundation Model · Robust Speech Recognition via Large-Scale Weak Supervision
Section: Speech and Text · multilingual ASR; the modern open default.
Section: 2. Model Codebases & Model Families · Canonical open speech-to-text model codebase with widespread ecosystem support and many downstream implementations.
Section: Other · Robust Speech Recognition via Large-Scale Weak Supervision
Section: Official
Section: Artificial Intelligence · Whisper is a general-purpose speech recognition model that can be run locally offline. It can transcribe audio from and to multiple languages.
Section: AI and Agents · A general-purpose automatic speech recognition model trained on 680k hours of multilingual and multitask supervised data.
Langchain integrates various providers like Anthropic, AWS, and OpenAI, and offers tools for components such as LLMs, chat models, and data analysis, supporting functionalities from Alpha Vantage to YouTube github | docs
(MIT) provides modules for structured outputs at different levels of abstraction, including output parsers for text completion endpoints, Pydantic programs for mapping prompts to structured outputs using function calling or output parsing, and pre-defined Pydantic programs for specific output types.
robot: The free, Open Source alternative to OpenAI, Claude and others. Self-hosted and local-first. Drop-in replacement for OpenAI, running on consumer-grade hardware. No GPU required. Runs gguf, transformers, diffusers and many more models architectures. Features: Generate Text, Audio, Video,…
(from Hpcaitech) - A Unified Deep Learning System for Large-Scale Parallel Training (1D, 2D, 2.5D, 3D and sequence parallelism, and ZeRO protocol).
February 2026 release making human oversight a native workflow primitive: suspend execution at critical decision points, expose review-and-edit UI mid-flow, and route subsequent execution based on human action (approve/reject/escalate). Demonstrates how HITL transitions from bolt-on approval gates…
Mem0 is an intelligent memory layer for Large Language Models that enhances personalized AI experiences by retaining and utilizing contextual information across various applications. github | website | docs | discord | twitter | github profile | linkedin
Microsoft's multi-agent conversation framework with a complete AgentChat layer covering agent loop, tool integration, termination conditions, and human-in-the-loop. The most comprehensive open-source reference for large-scale multi-agent harness design.
Flowise simplifies the creation of applications leveraging large language models (LLMs) by providing a drag-and-drop interface for customizing AI workflows, offering easy installation, Docker support, development tools, and documentation for integrating various functionalities such as…