LM Studio
discover, download and run local LLMs
A curated list of awesome platforms, tools, practices and resources that helps run LLMs locally
This page lists names, links and short descriptions. The original list on GitHub is the source and belongs to its authors.
unified web UI for training and running open models like Qwen, DeepSeek, and Gemma locally
an open source alternative to ChatGPT that runs 100% offline on your computer
user-friendly desktop client app for AI models/LLMs
a local LLM server with GPU and NPU Acceleration
a fast serving framework for large language models and vision language models
LLM inference server with continuous batching & SSD caching for Apple Silicon — managed from the macOS menu bar
provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs
a datacenter scale distributed inference serving framework
generate text and fine-tune large language models on Apple silicon with MLX
fast, flexible LLM inference
kernel library for LLM serving
Google's production-ready, high-performance, open-source inference framework for deploying Large Language Models on edge devices
a package for inference and fine-tuning of Vision Language Models (VLMs) on your Mac using MLX
a lightweight yet high-performance inference framework for Large Language Models
on-device AI across mobile, embedded and edge for PyTorch
Google's on-device framework for high-performance ML & GenAI deployment on edge platforms, via efficient conversion, runtime, and optimization
connect home devices into a powerful cluster to accelerate LLM inference
llama.cpp fork with additional SOTA quants and improved performance
a speed-of-light LLM inference engine
large-scale LLM inference engine based on vLLM
run LLMs on AMD Ryzen™ AI NPUs
a Hybrid LLM runtime which focuses on efficient running of larger models on consumer grade VRAM limited hardware
run LLMs on Intel Arc™ Pro B60 and B70 GPUs
vLLM for AMD gfx906 GPUs, e.g. Radeon VII / MI50 / MI60
User-friendly AI Interface (Supports Ollama, OpenAI API, ...)
LLM UI with advanced features, easy setup, and multiple backend support
LLM Frontend for Power Users
Use your locally running AI models to assist you in your web browsing
benchmark & compare the best AI models
understand the AI landscape to choose the best model and provider for your use case
a continuously evolving and decontaminated benchmark for software engineering LLMs
measure whether AI models challenge nonsensical prompts instead of confidently answering them
explore list of the open-source LLM models
small-scale manual performance comparison benchmark
a list sorted by size (on disk) for each score
evaluating AI agents' real-world cybersecurity capabilities at scale
a benchmark for evaluating multi-hop, multi-source tool-calling in AI agents
a benchmark for evaluating multi-hop, multi-source tool-calling in AI agents
powered by Alibaba Cloud
a pioneering French artificial intelligence startup
a profile of a Chinese multinational technology conglomerate and holding company
focusing on making AI more accessible to everyone (GGUFs etc.)
providing GGUF versions of popular LLMs
a private non-profit organization engaged in AI research and development
a team of researchers and engineers curating the best open reasoning datasets
a collection of the DeepSeek V4 LLMs
a collection of the latest generation Qwen LLMs
a family of open models from NVIDIA with open weights, training data and recipes, delivering leading efficiency and accuracy for building specialized AI agents
a family of open models built by Google DeepMind, that are multimodal, handling text and image input (with audio supported on small models) and generating text output
The first flaship models from Mistral AI handling instruction-following, reasoning, and coding in a single set of opened-weights
a collection of open-weight models from OpenAI, designed for powerful reasoning, agentic tasks, and versatile developer use cases
a deployment-optimized large language model developed by NVIDIA, derived from OpenAI's gpt-oss-120b
a collection of Tencent's open-source efficient LLMs designed for versatile deployment across diverse computational environments
a family of small language, multi-modal and reasoning models from Microsoft
a collection of models from NVIDIA, trained on 5M reasoning traces for math, code and science
a collection of open-source, native multimodal agentic models from Moonshot AI that advances practical capabilities in long-horizon coding, coding-driven design, proactive autonomous execution, and swarm-based task orchestration
a Z.ai's flagship model for long-horizon tasks
a native hybrid reasoning model from inclusionAI, operating with 124B total and 5.1B active parameters
efficient language models from IBM for multilingual generation, coding, RAG, and AI assistant workflows
a collection of open-source models for agentic tasks and coding
LG's First Open-Weight Vision-Language Model for Industrial Intelligence
Mixture-of-Experts vision-language model for developers who need to scale agentic workflows that combine perception, search, and reasoning
a compact agentic model purpose-built for real-world co-work
a collection of SOTA on-device LLMs, small yet powerful
a collection of Qwen's open-weight language models designed specifically for coding agents and local development
a couple of agentic LLMs for software engineering tasks, excelling at using tools to explore codebases, edit multiple files, and power SWE Agents
an assistant model trained by JetBrain
a native multimodal model with 1M context
a collection of SOTA models for real-world dev & agents
a 118B total parameter Mixture-of-Experts model with 8B activated parameters per token, designed for agentic coding and long-horizon work
a family of code-search models from Microsoft powering the Explore subagent for coding agents
a 9-billion parameter coding agent model built by Tesslate, fine-tuned on top of Qwen3.5-9B's hybrid architecture
a competitive programming model post-trained on Qwen3-14B via reinforcement learning
a code model developed by Moore Threads for PyTorch-to-CUDA/MUSA native kernel generation
a collection of the natively end-to-end multilingual omni-modal foundation models from Qwen
a collection of open source multimodal models with native tool use from Zhipu AI
a unified text-to-image generation and image editing model in the Qwen family
a collection of the most powerful vision-language models in the Qwen series to date
an image generation model
a 6B text-to-image model for UI, infographics, posters, and other text-rich visual designs
multimodal models from IBM built for visual document analysis and image understanding
a collection of image generation models from Tencent
a collection of video generation models from Tencent
a collection of models for multimodal video understanding and creation
a collection of VLMs with efficient vision encoding from Apple
multimodal models with leading performance
a collection of vision-language models, designed for on-device deployment
a vision-language model (VLM) designed for video understanding at massive scale
a state-of-the-art model for automatic speech recognition (ASR) and speech translation from OpenAI
a collection of open, state-of-the-art, production‑ready enterprise speech models from NVIDIA for ASR, TTS, Speaker Diarization and S2SOpenAI
a 11B end-to-end, real-time speech full duplex (FD) model from NVIDIA for conversational AI that jointly performs streaming speech understanding and speech generation
a collection of models that support language identification and ASR for 52 languages and dialects
a collection of TTS models that cover 10 major languages as well as multiple dialectal voice profiles to meet global application needs
a collection of compact and efficient speech-language models from IBM, specifically designed for multilingual automatic speech recognition (ASR) and bidirectional automatic speech translation (AST)
an enhancement of Mistral Small 3, incorporating state-of-the-art audio input capabilities while retaining best-in-class text performance
a multilingual, realtime speech-transcription model and among the first open-source solutions to achieve accuracy comparable to offline systems with a delay of <500ms
frontier, open-weights text-to-speech model that’s fast, instantly adaptable, and produces lifelike speech for voice agents
first production-grade open-source TTS model
a collection of frontier text-to-speech models from Microsoft
a collection of open-source realistic text-to-speech models designed for lightweight deployment and high-quality voice synthesis
a streaming version of a novel end-to-end neural model for speaker diarization from NVIDIA
a set of tools to build retrieval-augmented generation (RAG) systems, improve search and ranking accuracy, and extract structured data from complex docs
a collection of the latest proprietary Qwen models, specifically designed for text embedding and ranking tasks
an addition to the Qwen embedding models, specifically designed for multimodal information retrieval and cross-modal understanding
a collection of the latest proprietary Qwen models, engineered to refine embedding results
a compact 3B-parameter, policy-adaptive multimodal safety classifier
a collection of safety models from IBM for detecting risks, toxicity, and hallucinations in LLM workflows
a collection of safety moderation models built upon Qwen3
a collection of models from NVIDIA for content safety, topic-following and security guardrails
a small language model (SLM) that uses Google's Gemma-3-4B-it as the base and is fine-tuned by NVIDIA on multimodal, multilingual, and reasoning-oriented content-safety datasets
a collection of lightweight text-sanitization models from NVIDIA designed to remove or abstract sensitive information from text according to a user-provided sanitization instruction
a collection of policy-adaptive multimodal LLM Guardrails with dynamic reasoning
a family of safety-aligned instruction models from Microsoft trained with HARC
a collection of safety reasoning models built-upon gpt-oss from OpenAI
a bidirectional token-classification model from OpenAI for personally identifiable information (PII) detection and masking in text
a safeguard model designed to detect and mitigate both safety risks and security threats in LLM interactions
a collection of multimodal foundation models for scientific intelligence and long-horizon agents
a collection of Visual Language Models for computer/mobile use, tool calls, and code
a suit of multilingual MoE models with highly-sparse architectures
a diffusion model built for generative UI
a state-of-the-art 8B orchestration model designed to solve complex, multi-turn agentic tasks by coordinating a diverse set of expert models and tools
the fastest LLM router model that aligns to subjective usage preferences
a collection of real-time interactive video world models
a collection of everything related (models, datasets etc.) to 3D assets generation from Tencent
a novel framework for high-dynamic interactive video generation in game environments
a model from Netflix that removes objects from videos along with all interactions they induce on the scene — not just secondary effects like shadows and reflections, but physical interactions like objects falling when a person is removed
hundreds of models & providers, one command to find what runs on your hardware
reliable model swapping for any local OpenAI compatible server - llama.cpp, vllm, etc.
super-fast structured outputs
a powerful platform that allows you to create, deploy, and manage continuous AI agents that automate complex workflows
AI agent toolkit: coding agent CLI, unified LLM API, TUI & web UI libraries, Slack bot, vLLM pods
the all-in-one Desktop & Docker AI application with built-in RAG, AI agents, No-code agent builder, MCP compatibility, and more
the leading framework for building LLM-powered agents over your data
a full-stack framework for building Multi-Agent Systems with memory, knowledge and reasoning
a lightweight, powerful framework for multi-agent workflows
run OpenClaw more securely inside NVIDIA OpenShell with managed inference
a Python agent framework designed to help you quickly, confidently, and painlessly build production grade applications and workflows with Generative AI
an open-source framework to build, manage and run useful Autonomous AI Agents
a framework for building, orchestrating and deploying AI agents and multi-agent workflows with support for Python and .NET
all-in-one open-source AI framework for semantic search, LLM orchestration and language model workflows
a high-performance proxy server that handles the low-level work in building agents: like applying guardrails, routing prompts to the right agent, and unifying access to LLMs, etc.
open-source framework for building AI-powered apps in JavaScript, Go, and Python, built and used in production by Google
privacy-first, fully local AI workspace with Ollama LLM chat, tool calling, agent builder, Stable Diffusion, and embedded n8n-style automation
an open-source library for efficiently connecting and optimizing teams of AI agents
building blocks for rapid development of GenAI applications
Playwright MCP server
GitHub's official MCP Server
Chrome DevTools for coding agents
a MCP for Claude Desktop / Claude Code / Windsurf / Cursor to build n8n workflows for you
AWS MCP Servers — helping you get the most out of AWS, wherever you use MCP
MCP server for Atlassian tools (Confluence, Jira)
zero-dependency, token-efficient database MCP server for Postgres, MySQL, SQL Server, MariaDB, SQLite
Python ETL framework for stream processing, real-time analytics, LLM pipelines and RAG
build real-time knowledge graphs for AI Agents
AI orchestration framework to build customizable, production-ready LLM applications, best suited for building RAG, question answering, semantic search or conversational agent chatbots
an open-source Python RAG framework for SQL generation and related functionality
make entire codebase the context for any coding agent
a fully extensible and explainable workplace AI platform for enterprise search and workflow automation
a AI coding agent built for the terminal
a next-generation code editor designed for high-performance collaboration with humans and AI
autonomous coding agent right in your IDE, capable of creating/editing files, executing commands, using the browser, and more with your permission every step of the way
an open-source, extensible AI agent that goes beyond code suggestions
create, share, and use custom AI code assistants with our open-source IDE extensions and hub of models, rules, prompts, docs, and other building blocks
an open-source GitHub Copilot alternative, set up your own LLM-powered code completion server
an open-source Cursor alternative, use AI agents on your codebase, checkpoint and visualize changes, and bring any model or host locally
a whole dev team of AI agents in your code editor
the best way to get AI coding agents to solve hard problems in complex codebases
Agentic Development Environment based on OpenCode AI agent
neovim AI agent done right
the leading open-source AI copilot for JetBrains
a natural language interface for computers
the Docker Container for Computer-Use AI Agents
a simple screen parsing tool towards pure vision based GUI agent
a framework to enable multimodal models to operate a computer
a browser-based desktop where AI Agent operates every app through natural language, from MiniMaxAI
turn entire websites into LLM-ready markdown or structured data
make websites accessible for AI agents
a framework for Web Testing and Automation
the AI Browser Automation Framework
open-source Chrome extension for AI-powered web automation
the highest-scoring AI memory system ever benchmarked
memory engine and app that is extremely fast, scalable
an open-source memory framework for AI companions
a memory mechanism for agents that learns from both successful and failed trajectories, with reasoning stored as memory content
an open-source LLM engineering platform: LLM Observability, metrics, evals, prompt management, playground, datasets. Integrates with OpenTelemetry, Langchain, OpenAI SDK, LiteLLM, and more
debug, evaluate, and monitor your LLM applications, RAG systems, and agentic workflows with comprehensive tracing, automated evaluations, and production-ready dashboards
an open-source observability for your LLM application, based on OpenTelemetry
an open-source evaluation & testing for AI & LLM systems
an open-source LLMOps platform: prompt playground, prompt management, LLM evaluation, and LLM observability all in one place
open-source library for scalable, reproducible evaluation of AI models and benchmarks
an open-source implementation of Notebook LM with more flexibility and features
an open-source alternative to Perplexity AI, the AI-powered search engine
an LLM based autonomous agent that conducts deep local and web research on any topic and generates a long report with citations
automate the most critical and valuable aspects of the industrial R&D process
fully local web research and report writing assistant
an AI-powered research assistant for deep, iterative research
an AI-powered research application designed to streamline complex research tasks
fully automatic censorship removal for language models
a Python library for using and training embedding and reranker models for applications like retrieval augmented generation, semantic search, and more
an easy-to-use, high-performance open-source RLHF framework built on Ray, vLLM, ZeRO-3 and HuggingFace Transformers, designed to make RLHF training simple and accessible
the easiest tool for fine-tuning LLM models, synthetic data generation, and collaborating on datasets
an enterprise-facing reinforcement learning framework for LLM and VLM post-training, forked from and co-evolving with slime
an interface library for RL post training with environments
scalable toolkit for efficient model reinforcement
train an open-source LLM on new facts
evaluate and improve models and agents using environments
train speculative decoding models effortlessly and port them smoothly to SGLang serving
an open-source toolkit from NVIDIA for easily adding programmable guardrails to LLM-based conversational systems
instant, concurrent, secure & lightweight sandbox for AI agents
open source DeepWiki: AI-powered wiki generator for GitHub/Gitlab/Bitbucket repositories
a personal, self-hosted web application designed for transcribing audio recordings
an open-source AI presentation generator and API
exploration to advanced multimodal generation
a powerful, self-hosted AI photo stylizer built for performance and privacy
local open-source micro-agents that observe, log and react, all while keeping your data private and secure
a powerful, open-source AI agent that controls your Android or IOS device using natural language
build AI applications that can see, hear, and speak using your screens, microphones, and cameras as inputs
a zero-dependency prompt manager/catalog/library in a single HTML file
tests of pcs, laptops, gpus etc. capable of running LLMs
reviews of various builds designed for LLM inference
practical and insightful tutorials on running LLMs locally
information about developing on NVIDIA Jetson Development Kits
tests of various types of hardware capable of running LLMs
estimate the RAM requirements of any GGUF model instantly
calculate how many GPUs you need to deploy LLMs
CUDA on non-NVIDIA GPUs
toolboxes for GenAI on AMD Ryzen AI MAX+: containerized environments for LLMs, Image Generation, and Fine-tuning
a website to gather important information and practical guides for systems powered by AMD Ryzen AI MAX and MAX+ processors
random AI notes for working with local models or playing around with random machine learning bits
a full-stack implementation of an LLM like ChatGPT in a single, clean, minimal, hackable, dependency-lite codebase, designed to run on a single 8XH100 node via scripts like speedrun.sh, that run the entire pipeline start to end
Docs for GGUF quantization (unofficial)
guides, papers, lecture, notebooks and resources for prompt engineering
a comprehensive collection of tutorials and implementations for Prompt Engineering techniques, ranging from fundamental concepts to advanced strategies
a quick-start handbook for effective prompts by Google
prompt engineering by Google
prompt engineering by Anthropic
Prompt Engineering Interactive Tutorial by Anthropic
a collection of system prompts extracted from AI tools
a collection of extracted System Prompts from popular chatbots like ChatGPT, Claude & Gemini
Prompt used to steer behavior of OpenAI's Codex
a frontier, first-principles handbook inspired by Karpathy and 3Blue1Brown for moving beyond prompt engineering to the wider discipline of context design, orchestration, and optimization
a comprehensive survey on Context Engineering: from prompt engineering to production-grade AI systems
vLLM’s reference system for K8S-native cluster-wide deployment with community-driven performance optimization
an agentic skills framework & software development methodology that works
tutorials and implementations for various Generative AI Agent techniques
a curated collection of AI agent use cases across various industries
principles for building reliable LLM applications
end-to-end, code-first tutorials covering every layer of production-grade GenAI agents, guiding you from spark to scale with proven patterns and reusable blueprints for real-world launches
a simple, open format for guiding coding agents
a simple, open format for giving agents new capabilities and expertise
Hugging Face Skills are definitions for AI/ML tasks like dataset creation, model training and evaluation
one-stop handbook for building, deploying, and understanding LLM agents with 60+ skeletons, tutorials, ecosystem guides, and evaluation tools
601 real-world gen AI use cases from the world's leading organizations by Google
a practical guide to building agents by OpenAI
ready-to-run cloud templates for RAG, AI pipelines, and enterprise search with live data
various advanced techniques for Retrieval-Augmented Generation (RAG) systems
an advanced Retrieval-Augmented Generation (RAG) solution for complex question answering that uses sophisticated graph based algorithm to handle the tasks
a collection of modular RAG techniques, implemented in LangChain + Python
everything jamesob knows about running LLMs locally
Go-to subreddit for local/open-source LLM topics.
monitoring updates & fresh releases related to LLMs, diffusion models and Generative AI
hesreallyhim/awesome-claude-code
A hand-picked collection of the finest of resources for the most awesome of agents, Claude Code, the undisputed champion of coding companions, from the unstoppable team…
VoltAgent/awesome-agent-skills
A curated collection of 1000+ agent skills from official dev teams and the community, compatible with Claude Code, Codex, Gemini CLI, Cursor, and more.
josephmisiti/awesome-machine-learning
A curated list of awesome Machine Learning frameworks, libraries and software.
EthicalML/awesome-production-machine-learning
A curated list of awesome open source libraries to deploy, monitor, version and scale your machine learning
academic/awesome-datascience
:memo: An awesome Data Science repository to learn and apply for real world problems.
analysis-tools-dev/static-analysis
⚙️ A curated list of static analysis (SAST) tools and linters for all programming languages, config files, build tools, and more. The focus is on tools which improve…