Awesome LLM Resources
🧑🚀 全世界最好的LLM资料总结(多模态生成、Agent、辅助编程、AI审稿、数据处理、模型训练、模型推理、o1 模型、MCP、小语言模型、视觉语言模型) | Summary of the world's best LLM resources.
This page lists names, links and short descriptions. The original list on GitHub is the source and belongs to its authors.
推荐 Suggestion
数据 Data
AotoLabel
Label, clean and enrich text datasets with LLMs.
LabelLLM
The Open-Source Data Annotation Platform.
data-juicer
A one-stop data processing system to make data higher-quality, juicier, and more digestible for LLMs!
OmniParser
a native Golang ETL streaming parser and transform library for CSV, JSON, XML, EDI, text, etc.
MinerU (🔥)
MinerU is a one-stop, open-source, high-quality data extraction tool, supports PDF/webpage/e-book extraction.
PDF-Extract-Kit
A Comprehensive Toolkit for High-Quality PDF Content Extraction.
Parsera
Lightweight library for scraping web-sites with LLMs.
Sparrow
Sparrow is an innovative open-source solution for efficient data extraction and processing from various documents and images.
Docling
Get your documents ready for gen AI.
GOT-OCR2.0
OCR Model.
LLM Decontaminator
Rethinking Benchmark and Contamination for Language Models with Rephrased Samples.
DataTrove
DataTrove is a library to process, filter and deduplicate text data at a very large scale.
llm-swarm
Generate large synthetic datasets like Cosmopedia.
Distilabel
Distilabel is a framework for synthetic data and AI feedback for engineers who need fast, reliable and scalable pipelines based on verified research papers.
Common-Crawl-Pipeline-Creator
The Common Crawl Pipeline Creator.
Tabled
Detect and extract tables to markdown and csv.
Zerox
Zero shot pdf OCR with gpt-4o-mini.
DocLayout-YOLO
Enhancing Document Layout Analysis through Diverse Synthetic Data and Global-to-Local Adaptive Perception.
TensorZero
make LLMs improve through experience.
Promptwright
Generate large synthetic data using a local LLM.
pdf-extract-api
Document (PDF) extraction and parse API using state of the art modern OCRs + Ollama supported models.
pdf2htmlEX
Convert PDF to HTML without losing text or format.
Extractous
Fast and efficient unstructured data extraction. Written in Rust with bindings for many languages.
MegaParse
File Parser optimised for LLM Ingestion with no loss.
MarkItDown
Python tool for converting files and office documents to Markdown.
datasketch
datasketch gives you probabilistic data structures that can process and search very large amount of data super fast, with little loss of accuracy.
semhash
lightweight and flexible tool for deduplicating datasets using semantic similarity.
ReaderLM-v2
a 1.5B parameter language model that converts raw HTML into beautifully formatted markdown or JSON.
Bespoke Curator
Data Curation for Post-Training & Structured Data Extraction.
LangKit
An open-source toolkit for monitoring Large Language Models (LLMs). Extracts signals from prompts & responses, ensuring safety & security.
olmOCR
A toolkit for training language models to work with PDF documents in the wild.
Easy Dataset (🔥)
A powerful tool for creating fine-tuning datasets for LLM.
BabelDOC
PDF scientific paper translation and bilingual comparison library.
Dolphin
Document Image Parsing via Heterogeneous Anchor Prompting.
EasyDistill
Easy Knowledge Distillation for Large Language Models.
ContextGem
a free, open-source LLM framework that makes it radically easier to extract structured data and insights from documents.
OCRFlux
a lightweight yet powerful multimodal toolkit that significantly advances PDF-to-Markdown conversion, excelling in complex layout handling, complicated table parsing and cross-page content merging.
DataFlow
Easy Data Preparation with latest LLMs-based Operators and Pipelines.
DatasetLoom (multimodal)
一个面向多模态大模型训练的智能数据集构建与评估平台.
Chandra
a highly accurate OCR model that converts images and PDFs into structured HTML/Markdown/JSON while preserving layout information.
HunyuanOCR
a leading end-to-end OCR expert VLM powered by Hunyuan's native multimodal architecture.
DeepSeek-OCR-2
Visual Causal Flow.
PaddleOCR-VL-1.5 (🔥)
Towards a Multi-Task 0.9B VLM for Robust In-the-Wild Document Parsing.
GLM-OCR
a multimodal OCR model for complex document understanding, built on the GLM-V encoder–decoder architecture.
微调 Fine-Tuning
LLaMA-Factory (🔥)
Unify Efficient Fine-Tuning of 100+ LLMs.
360-LLaMA-Factory
Unify Efficient Fine-Tuning of 100+ LLMs. (add Sequence Parallelism for supporting long context training)
unsloth (🔥)
2-5X faster 80% less memory LLM finetuning.
TRL
Transformer Reinforcement Learning.
Firefly
Firefly: 大模型训练工具,支持训练数十种大模型
Xtuner
An efficient, flexible and full-featured toolkit for fine-tuning large models.
torchtune
A Native-PyTorch Library for LLM Fine-tuning.
Swift
Use PEFT or Full-parameter to finetune 200+ LLMs or 15+ MLLMs.
AutoTrain
A new way to automatically train, evaluate and deploy state-of-the-art Machine Learning models.
OpenRLHF
An Easy-to-use, Scalable and High-performance RLHF Framework (Support 70B+ full tuning & LoRA & Mixtral & KTO).
Ludwig
Low-code framework for building custom LLMs, neural networks, and other AI models.
mistral-finetune
A light-weight codebase that enables memory-efficient and performant finetuning of Mistral's models.
aikit
Fine-tune, build, and deploy open-source LLMs easily!
H2O-LLMStudio
H2O LLM Studio - a framework and no-code GUI for fine-tuning LLMs.
LitGPT
Pretrain, finetune, deploy 20+ LLMs on your own data. Uses state-of-the-art techniques: flash attention, FSDP, 4-bit, LoRA, and more.
LLMBox
A comprehensive library for implementing LLMs, including a unified training pipeline and comprehensive model evaluation.
workbench-llamafactory
This is an NVIDIA AI Workbench example project that demonstrates an end-to-end model development workflow using Llamafactory.
TinyLLaVA Factory
A Framework of Small-scale Large Multimodal Models.
LLM-Foundry
LLM training code for Databricks foundation models.
lmms-finetune
A unified codebase for finetuning (full, lora) large multimodal models, supporting llava-1.5, qwen-vl, llava-interleave, llava-next-video, phi3-v etc.
Simplifine
Simplifine lets you invoke LLM finetuning with just one line of code using any Hugging Face dataset or model.
Transformer Lab
Open Source Application for Advanced LLM Engineering: interact, train, fine-tune, and evaluate large language models on your own computer.
Liger-Kernel
Efficient Triton Kernels for LLM Training.
ChatLearn
A flexible and efficient training framework for large-scale alignment.
nanotron
Minimalistic large language model 3D-parallelism training.
Proxy Tuning
Tuning Language Models by Proxy.
Effective LLM Alignment
Effective LLM Alignment Toolkit.
Autotrain-advanced
AutoTrain Advanced is a no-code solution that allows you to train machine learning models in just a few clicks.
Meta Lingua
a lean, efficient, and easy-to-hack codebase to research LLMs.
Vision-LLM Alignemnt
This repository contains the code for SFT, RLHF, and DPO, designed for vision-based LLMs, including the LLaVA models and the LLaMA-3.2-vision models.
finetune-Qwen2-VL
Quick Start for Fine-tuning or continue pre-train Qwen2-VL Model.
Online-RLHF
A recipe for online RLHF and online iterative DPO.
InternEvo
an open-sourced lightweight training framework aims to support model pre-training without the need for extensive dependencies.
veRL (🔥)
Volcano Engine Reinforcement Learning for LLM.
Axolotl
Axolotl is designed to work with YAML config files that contain everything you need to preprocess a dataset, train or fine-tune a model, run model inference or evaluation, and much more.
Oumi
Everything you need to build state-of-the-art foundation models, end-to-end.
Kiln
The easiest tool for fine-tuning LLM models, synthetic data generation, and collaborating on datasets.
DeepSeek-671B-SFT-Guide
An open-source solution for full parameter fine-tuning of DeepSeek-V3/R1 671B, including complete code and scripts from training to inference, as well as some practical experiences and conclusions.
MLX-VLM
MLX-VLM is a package for inference and fine-tuning of Vision Language Models (VLMs) on your Mac using MLX.
RL-Factory
Train your Agent model via our easy and efficient framework.
RM-Gallery
A One-Stop Reward Model Platform.
ART
rain multi-step agents for real-world tasks using GRPO. Give your agents on-the-job training.
LMMs-Engine
A simple, any-to-any modality framework for pretraining and finetuning. Lean, flexible, and built for research.
dLLM
a library that unifies the training and evaluation of diffusion language models, bringing transparency and reproducibility to the entire development pipeline. diffusion
Miles
an enterprise-facing reinforcement learning framework for large-scale MoE post-training and production workloads.
Skills
a collection of pipelines to improve "skills" of large language models (LLMs).
Twinkle
a lightweight, client-server training framework engineered with modular, high-cohesion interfaces.
NeMo AutoModel
Pytorch Distributed native training library for LLMs/VLMs with OOTB Hugging Face support.
VeOmni
Scaling Any Modality Model Training with Model-Centric Distributed Recipe Zoo.
Soup
One-config CLI for LLM post-training (SFT/DPO/GRPO/KTO/ORPO). Layer streaming trains an 8B model on a 4 GB laptop GPU by streaming the frozen base from host RAM one decoder layer at a time.
Agentic RL
veRL (🔥)
Volcano Engine Reinforcement Learning for LLM.
AReaL:
AntGroup/Tsinghua
slime (🔥):
LLM post-training framework for RL Scaling from THUDM. Supports SFT and RL training with multi-turn compilation feedback, powering projects like TritonForge for automated GPU kernel generation. Apache 2.0 licensed.
Agent Lightning:
The absolute trainer to light up AI agents.
Molt:
NVIDIA (NeMo Labs)
prime-rl:
Agentic RL Training at Scale from Prime Intellect. Framework for large-scale reinforcement learning capable of scaling to 1000+ GPUs with fully asynchronous RL, FSDP2 training, and vLLM inference. Apache 2.0 licensed.
推理 Inference
ollama (🔥)
Get up and running with Llama 3, Mistral, Gemma, and other large language models.
Open WebUI
User-friendly WebUI for LLMs (Formerly Ollama WebUI).
Text Generation WebUI
A Gradio web UI for Large Language Models. Supports transformers, GPTQ, AWQ, EXL2, llama.cpp (GGUF), Llama models.
Xinference
A powerful and versatile library designed to serve language, speech recognition, and multimodal models.
LlamaIndex
A data framework for your LLM applications.
lobe-chat
an open-source, modern-design LLMs/AI chat framework. Supports Multi AI Providers, Multi-Modals (Vision/TTS) and plugin system.
TensorRT-LLM
TensorRT-LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and build TensorRT engines that contain state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs.
vllm (🔥)
A high-throughput and memory-efficient inference and serving engine for LLMs.
LlamaChat
Chat with your favourite LLaMA models in a native macOS app.
NVIDIA ChatRTX
ChatRTX is a demo app that lets you personalize a GPT large language model (LLM) connected to your own content—docs, notes, or other data.
LM Studio (🔥)
Discover, download, and run local LLMs.
chat-with-mlx
Chat with your data natively on Apple Silicon using MLX Framework.
LLM Pricing
Quickly Find the Perfect Large Language Models (LLM) API for Your Budget! Use Our Free Tool for Instant Access to the Latest Prices from Top Providers.
Open Interpreter
A natural language interface for computers.
Chat-ollama
An open source chatbot based on LLMs. It supports a wide range of language models, and knowledge base management.
MemGPT
Create LLM agents with long-term memory and custom tools.
koboldcpp
A simple one-file way to run various GGML and GGUF models with KoboldAI's UI.
LLMFarm
llama and other large language models on iOS and MacOS offline using GGML library.
enchanted
Enchanted is iOS and macOS app for chatting with private self hosted language models such as Llama2, Mistral or Vicuna using Ollama.
Jan
Jan is an open source alternative to ChatGPT that runs 100% offline on your computer. Multiple engine support (llama.cpp, TensorRT-LLM).
LMDeploy
LMDeploy is a toolkit for compressing, deploying, and serving LLMs.
RouteLLM
A framework for serving and evaluating LLM routers - save LLM costs without compromising quality!
MInference
About To speed up Long-context LLMs' inference, approximate and dynamic sparse calculate the attention, which reduces inference latency by up to 10x for pre-filling on an A100 while maintaining accuracy.
SGLang (🔥)
SGLang is yet another fast serving framework for large language models and vision language models.
AirLLM
AirLLM optimizes inference memory usage, allowing 70B large language models to run inference on a single 4GB GPU card without quantization, distillation and pruning. And you can run 405B Llama3.1 on 8GB vram now.
LLMHub
LLMHub is a lightweight management platform designed to streamline the operation and interaction with various language models (LLMs).
LiteLLM (🔥)
Call all LLM APIs using the OpenAI format [Bedrock, Huggingface, VertexAI, TogetherAI, Azure, OpenAI, Groq etc.]
GuideLLM
GuideLLM is a powerful tool for evaluating and optimizing the deployment of large language models (LLMs).
LLM-Engines
A unified inference engine for large language models (LLMs) including open-source models (VLLM, SGLang, Together) and commercial models (OpenAI, Mistral, Claude).
OARC
ollama_agent_roll_cage (OARC) is a local python agent fusing ollama llm's with Coqui-TTS speech models, Keras classifiers, Llava vision, Whisper recognition, and more to create a unified chatbot agent for local, custom automation.
g1
Using Llama-3.1 70b on Groq to create o1-like reasoning chains.
MemoryScope
MemoryScope provides LLM chatbots with powerful and flexible long-term memory capabilities, offering a framework for building such abilities.
OpenLLM
Run any open-source LLMs, such as Llama 3.1, Gemma, as OpenAI compatible API endpoint in the cloud.
Infinity
The AI-native database built for LLM applications, providing incredibly fast hybrid search of dense embedding, sparse embedding, tensor and full-text.
optillm
an OpenAI API compatible optimizing inference proxy which implements several state-of-the-art techniques that can improve the accuracy and performance of LLMs.
LLaMA Box
LLM inference server implementation based on llama.cpp.
ZhiLight
A highly optimized inference acceleration engine for Llama and its variants.
DashInfer
DashInfer is a native LLM inference engine aiming to deliver industry-leading performance atop various hardware architectures.
LocalAI
The free, Open Source alternative to OpenAI, Claude and others. Self-hosted and local-first. Drop-in replacement for OpenAI, running on consumer-grade hardware. No GPU required.
ktransformers
A Flexible Framework for Experiencing Cutting-edge LLM Inference Optimizations.
SkyPilot
Run AI and batch jobs on any infra (Kubernetes or 14+ clouds). Get unified execution, cost savings, and high GPU availability via a simple interface.
Chitu
High-performance inference framework for large language models, focusing on efficiency, flexibility, and availability.
TokenSwift
From Hours to Minutes: Lossless Acceleration of Ultra Long Sequence Generation.
Cherry Studio (🔥)
a desktop client that supports for multiple LLM providers, available on Windows, Mac and Linux.
Shimmy
Python-free Rust inference server — OpenAI-API compatible. GGUF + SafeTensors, hot model swap, auto-discovery, single binary.
LlamaBarn
Run local LLMs on your Mac with a simple menu bar app.
Parallax
a distributed model serving framework that lets you build your own AI cluster anywhere.
xLLM
A high-performance inference engine for LLMs, optimized for diverse AI accelerators.
Rapid-MLX
OpenAI-compatible local LLM inference server for Apple Silicon, 2-4x faster than Ollama.
TokenSpeed
a speed-of-light LLM inference engine designed for agentic workloads, with TensorRT-LLM-level performance and vLLM-level usability. Our goal is to be the most performant inference engine for production agentic workloads.
FreeToken
Unlock datacenter-class intelligence on the hardware you already own.
评估 Evaluation
lm-evaluation-harness
A framework for few-shot evaluation of language models.
opencompass (🔥)
OpenCompass is an LLM evaluation platform, supporting a wide range of models (Llama3, Mistral, InternLM2,GPT-4,LLaMa2, Qwen,GLM, Claude, etc) over 100+ datasets.
llm-comparator
LLM Comparator is an interactive data visualization tool for evaluating and analyzing LLM responses side-by-side, developed.
EvalScope (🔥)
Streamlined and customizable framework for efficient large model (LLM, VLM, AIGC) evaluation and performance benchmarking. One-stop evaluation solution with 80+ benchmarks. Apache 2.0 licensed.
Weave
A lightweight toolkit for tracking and evaluating LLM applications.
MixEval
Deriving Wisdom of the Crowd from LLM Benchmark Mixtures.
Evaluation guidebook
If you've ever wondered how to make sure an LLM performs well on your specific task, this guide is for you!
Ollama Benchmark
LLM Benchmark for Throughput via Ollama (Local LLMs).
VLMEvalKit
Open-source evaluation toolkit of large vision-language models (LVLMs), support ~100 VLMs, 40+ benchmarks.
DeepEval
a simple-to-use, open-source LLM evaluation framework, for evaluating and testing large-language model systems.
Lighteval
Lighteval is your all-in-one toolkit for evaluating LLMs across multiple backends.
QwQ/eval
QwQ is the reasoning model series developed by Qwen team, Alibaba Cloud.
Evalchemy
A unified and easy-to-use toolkit for evaluating post-trained language models.
MathArena
Evaluation of LLMs on latest math competitions.
YourBench
A Dynamic Benchmark Generation Framework.
MedEvalKit
A Unified Medical Evaluation Framework.
OpenJudge
A Unified Framework for Holistic Evaluation and Quality Rewards.
体验 Usage
知识库 RAG
AnythingLLM
The all-in-one AI app for any LLM with full RAG and AI Agent capabilites.
RAGFlow
An open-source RAG (Retrieval-Augmented Generation) engine based on deep document understanding.
Dify
An open-source LLM app development platform. Dify's intuitive interface combines AI workflow, RAG pipeline, agent capabilities, model management, observability features and more, letting you quickly go from prototype to production.
FastGPT
A knowledge-based platform built on the LLM, offers out-of-the-box data processing and model invocation capabilities, allows for workflow orchestration through Flow visualization.
Langchain-Chatchat
基于 Langchain 与 ChatGLM 等不同大语言模型的本地知识库问答
QAnything
Question and Answer based on Anything.
Quivr
A personal productivity assistant (RAG) ⚡️🤖 Chat with your docs (PDF, CSV, ...) & apps using Langchain, GPT 3.5 / 4 turbo, Private, Anthropic, VertexAI, Ollama, LLMs, Groq that you can share with users ! Local & Private alternative to OpenAI GPTs & ChatGPT powered by retrieval-augmented generation.
RAG-GPT
RAG-GPT, leveraging LLM and RAG technology, learns from user-customized knowledge bases to provide contextually relevant answers for a wide range of queries, ensuring rapid and accurate information retrieval.
Verba
Retrieval Augmented Generation (RAG) chatbot powered by Weaviate.
FlashRAG
A Python Toolkit for Efficient RAG Research.
LightRAG
LightRAG helps developers with both building and optimizing Retriever-Agent-Generator pipelines.
GraphRAG-Ollama-UI
GraphRAG using Ollama with Gradio UI and Extra Features.
nano-GraphRAG
A simple, easy-to-hack GraphRAG implementation.
RAG Techniques
This repository showcases various advanced techniques for Retrieval-Augmented Generation (RAG) systems. RAG systems combine information retrieval with generative models to provide accurate and contextually rich responses.
ragas
Evaluation framework for your Retrieval Augmented Generation (RAG) pipelines.
kotaemon
An open-source clean & customizable RAG UI for chatting with your documents. Built with both end users and developers in mind.
RAGapp
The easiest way to use Agentic RAG in any enterprise.
TurboRAG
Accelerating Retrieval-Augmented Generation with Precomputed KV Caches for Chunked Text.
TEN
the Next-Gen AI-Agent Framework, the world's first truly real-time multimodal AI agent framework.
AutoRAG
RAG AutoML tool for automatically finding an optimal RAG pipeline for your data.
KAG
KAG is a knowledge-enhanced generation framework based on OpenSPG engine, which is used to build knowledge-enhanced rigorous decision-making and information retrieval knowledge services.
Fast-GraphRAG
RAG that intelligently adapts to your use case, data, and queries.
DB-GPT GraphRAG
DB-GPT GraphRAG integrates both triplet-based knowledge graphs and document structure graphs while leveraging community and document retrieval mechanisms to enhance RAG capabilities, achieving comparable performance while consuming only 50% of the tokens required by Microsoft's GraphRAG. Refer to…
Chonkie
The no-nonsense RAG chunking library that's lightweight, lightning-fast, and ready to CHONK your texts.
RAGLite
RAGLite is a Python toolkit for Retrieval-Augmented Generation (RAG) with PostgreSQL or SQLite.
CAG
CAG leverages the extended context windows of modern large language models (LLMs) by preloading all relevant resources into the model’s context and caching its runtime parameters.
MiniRAG
an extremely simple retrieval-augmented generation framework that enables small models to achieve good RAG performance through heterogeneous graph indexing and lightweight topology-enhanced retrieval.
XRAG
a benchmarking framework designed to evaluate the foundational components of advanced Retrieval-Augmented Generation (RAG) systems.
Rankify
A Comprehensive Python Toolkit for Retrieval, Re-Ranking, and Retrieval-Augmented Generation.
RAG-Anything
All-in-One RAG System.
智能体 Agents
AutoGen
AutoGen is a framework that enables the development of LLM applications using multiple agents that can converse with each other to solve tasks. AutoGen AIStudio
CrewAI
Framework for orchestrating role-playing, autonomous AI agents. By fostering collaborative intelligence, CrewAI empowers agents to work together seamlessly, tackling complex tasks.
Coze
ByteDance agent builder. Visual workflow. Plugin marketplace.
MobileAgent
The Powerful Mobile Device Operation Assistant Family.
Lagent
A lightweight framework for building LLM-based agents.
Qwen-Agent
Agent framework and applications built upon Qwen2, featuring Function Calling, Code Interpreter, RAG, and Chrome extension.
LinkAI
一站式 AI 智能体搭建平台
agentUniverse
agentUniverse is a LLM multi-agent framework that allows developers to easily build multi-agent applications. Furthermore, through the community, they can exchange and share practices of patterns across different domains.
LazyLLM
低代码构建多Agent大模型应用的开发工具
AgentScope
Start building LLM-empowered multi-agent applications in an easier way.
AgentField
Open-source control plane for building and operating AI agents like APIs at scale, with routing, memory, observability, identity, auth, and policy controls.
MoA
Mixture of Agents (MoA) is a novel approach that leverages the collective strengths of multiple LLMs to enhance performance, achieving state-of-the-art results.
Agently
AI Agent Application Development Framework.
OmAgent
A multimodal agent framework for solving complex tasks.
Tribe
No code tool to rapidly build and coordinate multi-agent teams.
CAMEL
First LLM multi-agent framework and an open-source community dedicated to finding the scaling law of agents.
PraisonAI
PraisonAI application combines AutoGen and CrewAI or similar frameworks into a low-code solution for building and managing multi-agent LLM systems, focusing on simplicity, customisation, and efficient human-agent collaboration.
IoA
An open-source framework for collaborative AI agents, enabling diverse, distributed agents to team up and tackle complex tasks through internet-like connectivity.
llama-agentic-system
Agentic components of the Llama Stack APIs.
Agent Zero
Agent Zero is not a predefined agentic framework. It is designed to be dynamic, organically growing, and learning as you use it.
Agents
An Open-source Framework for Data-centric, Self-evolving Autonomous Language Agents.
FastAgency
The fastest way to bring multi-agent workflows to production.
Swarm
Framework for building, orchestrating and deploying multi-agent systems. Managed by OpenAI Solutions team. Experimental framework.
PydanticAI
Agent Framework / shim to use Pydantic with LLMs.
Agentarium
open-source framework for creating and managing simulations populated with AI-powered agents.
smolagents
a barebones library for agents. Agents write python code to call tools and orchestrate other agents.
Cooragent
Cooragent is an AI agent collaboration community.
Agno
Agno is a lightweight library for building Agents with memory, knowledge, tools and reasoning.
Suna
Open Source Generalist AI Agent.
EvoAgentX
Building a Self-Evolving Ecosystem of AI Agents.
ii-agent
a new open-source framework to build and deploy intelligent agents.
OWL
Optimized Workforce Learning for General Multi-Agent Assistance in Real-World Task Automation.
OpenManus
No fortress, purely open ground. OpenManus is Coming.
JoyAgent-JDGenie
业界首个开源高完成度轻量化通用多智能体产品.
coze-studio
An AI agent development platform with all-in-one visual tools, simplifying agent creation, debugging, and deployment like never before.
OxyGent
An advanced Python framework that empowers developers to quickly build production-ready intelligent systems.
LazyCraft
LazyCraft 是一个基于 LazyLLM 构建的 AI Agent 应用开发与管理平台,旨在协助开发者以 低门槛、低成本 快速构建和发布大模型应用。
OpenAgents
AI Agent Networks for Open Collaboration.
SandBox
All-in-One Sandbox for AI Agents that combines Browser, Shell, File, MCP and VSCode Server in a single Docker container.
DeepAnalyze
First agentic LLM for autonomous data science, supporting specific data tasks (data preparation, analysis, modeling, visualization, and insight) and data-oriented deep research (produce analyst-grade research reports).
Astron Agent
Enterprise-grade, commercial-friendly agentic workflow platform for building next-generation SuperAgents.
Youtu-Agent
A simple yet powerful agent framework that delivers with open-source models.
MiroThinker
an open-source search agent model, built for tool-augmented reasoning and real-world information seeking, aiming to match the deep research experience of OpenAI Deep Research and Gemini Deep Research.
Nexent
A zero-code platform for auto-generating agents — no orchestration, no complex drag-and-drop required, using pure language to develop any agent you want.
Yunjue-Agent
A Fully Reproducible, Zero-Start In-Situ Self-Evolving Agent System for Open-Ended Tasks.
Hindsight
State-of-the-art long-term memory for AI agents by Vectorize. Open-source, self-hostable, with integrations for LangChain, CrewAI, LlamaIndex, MCP, and more.
AgentsMesh
The AI Agent Workforce Platform. Self-hostable multi-agent orchestration with remote AI workstations (AgentPods), PTY sandbox + git worktree isolation, channels-based agent collaboration, built-in Kanban, and per-pod MCP server. Supports Claude Code, Codex CLI, Gemini CLI, Aider, OpenCode.
BitFun
Open-source agentic development environment with a Rust/Tauri desktop app and CLI for coding, research, office work, browser and desktop automation, extensible through MCP, Skills, and custom agents.
DeepSeek Harness
Everything is a Plugin.
OpenSquilla
a token-efficient, microkernel AI agent.
PenguinHarness
Your Automated Agent Builder, Right on Your Desktop / Server.
FrontierAgent
an open-source agent runtime, terminal product, and evaluation suite for long-horizon research and file-based work.
研究 Research
PaperDebugger:
Paper Debugger is the best overleaf companion
claude-prism:
Offline-first scientific writing workspace powered by Claude, integrating LaTeX, Python, and 100+ scientific skills with local execution, Zotero integration, and privacy-focused design (2026)
PPTAgent:
Beyond text-to-slides generation with PPTEval multi-dimensional evaluation (EMNLP 2025)
PPT Master:
AI turns documents or topics into real, native PowerPoint decks—with native shapes, transitions and animations, data-backed charts and tables on demand, audio narration from speaker notes, and support for your own .pptx templates. · by Hugo He
Kami:
Kami是一款专注于将Markdown内容转化为高质量PDF排版工具的开源项目,它精准解决了创作者在使用常规编辑器导出文档时遭遇的格式错乱、字体生硬及页眉页脚配置繁琐等痛点。该项目在同类工具中展现出三大核心优势:内置经过精细调校的学术级排版引擎,能够自动处理行距、段首缩进与代码块样式,彻底免除了手动调整的冗余;提供可视化主题编辑界面,用户无需编写CSS即可实时预览并切换封面、页码及版权信息,大幅降低了定制门槛;深度集成Pandoc底层转换能力却剥离了复杂的命令行依赖,通过轻量级Web服务实现一键导出,在兼容性与易用性之间取得了卓越平衡。从技术架构来看,Kami本质上扮演了“数字排版工厂”的角…
Paper2Video:
First benchmark for automatic video generation from scientific papers (NeurIPS 2025)
Paper2Poster:
Multi-agent system with Parser-Planner-Painter architecture converting paper.pdf to editable poster.pptx, outperforms GPT-4o with 87% fewer tokens
AutoPR:
Fix issues with AI-generated pull requests, powered by ChatGPT
Paper2All:
AI-powered pipeline converting papers into interactive websites, posters, and multimedia presentations with "Let's Make Your Paper Alive!" philosophy
PaperBanana:
Automated academic illustration generation for AI scientists, converting research papers into publication-ready figures using VLMs and diffusion models with iterative refinement (PKU & Google Research, 6.2K+ stars, 2026)
EvoScientist:
Self-evolving AI scientist with 6 specialized sub-agents (plan/research/code/debug/analyze/write) and persistent memory, #1 on DeepResearch Bench II and AstaBench, supporting multi-provider LLMs and multi-channel deployment (Apache 2.0, 2026)
Auto-claude-code-research-in-sleep:
ARIS ⚔️ (Auto-Research-In-Sleep) — Lightweight Markdown-only skills for autonomous ML research: cross-model review loops, idea discovery, and experiment automation. No framework, no lock-in — works with Claude Code, Codex, OpenClaw, or any LLM agent.
Dr.Claw:
Open-source research workspace with sequential idea-to-paper pipelines and integrated autoresearch tool packs.
AutoResearchClaw:
April 2026 open-source human-in-the-loop system with six intervention modes (full-auto, gate-only, checkpoint, step-by-step, co-pilot, custom), SmartPause confidence-driven dynamic suspension, and Intervention Learning from human corrections. The cost-guardrail system — aborting runs that exceed…
NanoResearch:
End-to-end autonomous AI research engine that turns an idea into a complete LaTeX paper by dispatching real computational experiments to local GPUs or SLURM clusters, collecting actual results, generating figures/tables, and writing a data-grounded manuscript rather than LLM hallucinations…
Claude-scholar:
Semi-automated research assistant for academic research and software development, supporting Claude Code, Codex CLI, Kimi Code CLI, and OpenCode across ideation, coding, experiments, writing, and publication (Galaxy-Dawn, 4.5K+ stars, MIT License, 2026)
claude-scientific-skills:
by K-Dense - "A set of ready-to-use Agent Skills for research, science, engineering, analysis, finance and writing." That's their description - modest, simple. That's how you can tell this is really one of the best skills repos on GitHub. If you've ever thought about getting a PhD... just read all…
K-Dense BYOK:
Free, open-source desktop AI research assistant that runs locally and turns natural-language requests into real data analysis, literature search, figure generation, and manuscript review; ships with 149 scientific skills, 326 workflow templates, and 229 databases across genomics, proteomics, drug…
AutoResearch :
Andrej Karpathy's autonomous LLM research framework: AI agent runs overnight experiments on a real training setup, auto-editing code→5min training→evaluation in a loop, ~100 experiments per night on a single GPU
RD-Agent :
Open-source LLM-powered R&D agent framework automating data-driven AI solution building through automated research, development, and evolution; achieves top open-source performance on MLE-Bench with dual Researcher-Developer agents and supports research copilot, data mining, Kaggle, and quant R&D…
DeepScientist :
First system progressively surpassing human SOTA on frontier AI tasks (183.7%, 1.9%, 7.9% improvements), month-long autonomous discovery with 20,000+ GPU hours
academic-research-skills:
Comprehensive Claude Code skill suite covering the full academic pipeline from deep research and paper writing to multi-perspective peer review, revision, and finalization; features multi-agent teams, PRISMA systematic review, style calibration, claim-level citation audits, integrity gates, and…
代码 Coding
Cloi CLI
Local debugging agent that runs in your terminal.
Devin
Autonomous AI software engineer by Cognition with its own IDE, shell, browser, and cloud sandbox.
v0
Prompt-driven UI generation for React and Next.js, creating production-ready components.
Blot.new
Build, edit, and deploy full-stack web apps in the browser using natural language with one-click deployment.
cursor
AI Code Editor with Cloud Agents, JetBrains integration, and 30+ plugins from partners like Atlassian, Datadog, and GitLab.
Windsurf
agentic IDE, "where the work of developers and AI truly flow together, allowing for a coding experience that feels like literal magic"
cline
autonomous coding agent right in your IDE, capable of creating/editing files, executing commands, using the browser, and more with your permission every step of the way
Trae
Free AI IDE by ByteDance with Builder Mode, free access to GPT-4o, Claude Sonnet, and DeepSeek R1.
Roo Code
An AI-powered autonomous coding agent integrated directly into VS Code. #opensource
Kilo Code
Open Source AI coding assistant for planning, building, and fixing code. We're a superset of Roo, Cline, and our own features. Follow us: kilocode.ai/social
AugmentCode
AI coding platform with deep cross-repo codebase understanding via its Context Engine, built for large enterprise codebases.
Claude Code (🔥)
Claude Code is an agentic coding tool that lives in your terminal, understands your codebase, and helps you code faster by executing routine tasks, explaining complex code, and handling git workflows - all through natural language commands.
Gemini CLI
The official open-source AI agent that brings the power of Gemini directly into your terminal. Features context-aware coding assistance, file manipulation, and command execution capabilities.
Serena
Powerful MCP toolkit for coding agents providing semantic retrieval and editing capabilities. Integrates language servers for IDE-level code understanding. MIT licensed.
OpenCode
Open-source terminal AI agent (95K+ GitHub stars) supporting 75+ providers. Free, privacy-first, with LSP integration.
Kiro
Spec-driven development. Write specs → auto-generate tasks → implement. DevOps automation.
CodeX (🔥)
OpenAI's official autonomous coding agent CLI — the open-source reference implementation of the Codex harness with sandboxed tool execution, multi-file editing, and a streaming agent loop. Worth studying because it is the most widely adopted terminal-native coding agent harness and exposes the…
Kimi-CLI
[Archived] Legacy Python Kimi CLI, no longer maintained. Please use Kimi Code CLI: https://github.com/MoonshotAI/kimi-code (archived)
opencode
Open-source terminal-native AI coding agent with 131K+ stars and 2.5M+ monthly active developers. Provider-agnostic architecture supports 75+ LLM providers plus native LSP auto-configuration, multi-session parallel agents, and MCP extensibility. The build/plan agent split and client/server…
Multica
Managed agents platform where you assign tasks, track progress, and let agents compound skills between runs.
Atomic Agent
Local-first coding agent that runs open-weight models entirely on your machine via a llama.cpp fork. 56 built-in tools (browser, filesystem, git, memory, vision), MCP support, and a five-layer memory system. No account or API key required.
DeepSeek Harness
Everything is a Plugin.
视频 Video
HunyuanVideo
HunyuanVideo is an open-source video foundation model that demonstrates performance in video generation comparable to, or even surpassing, leading closed-source models. It encompasses a comprehensive framework integrating data curation, advanced architectural design, progressive model scaling and…
CogVideo
Tsinghua/Zhipu open-source, multiple sizes.See also: Software Reference → AI Video Generation Software
Open-Sora-Plan
This project aim to reproduce Sora (Open AI T2V model), we wish the open source community contribute to this project.
LTX-Video
LTX-Video is the first DiT-based video generation model that can generate high-quality videos in real-time. It can generate 24 FPS videos at 768x512 resolution, faster than it takes to watch them.
Step-Video-T2V
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model.
Step1X-Edit
Editing
Wan2.1-VACE
Editing
ICEdit
Editing
Wan2.1-FLF2V
首尾帧
MAGI-1
自回归模型
FramePack
Lets make video diffusion practical!
Wan2.2
Wan: Open and Advanced Large-Scale Video Generative Models.
MoGA
长视频
HunyuanVideo-1.5
HunyuanVideo-1.5: A leading lightweight video generation model.
Training
Official Python inference and LoRA trainer package for the LTX-2 audio–video generative model.
LongLive
LongLive: Real-time Interactive Long Video Generation.
(🔥)
High-Resolution Editable Toon Shading via Diffusion Models.
https://github.com/huggingface/diffusers
A library that provides pre-trained diffusion models for generating and editing images, audio, and video.
https://github.com/shengshu-ai/minWM
world model
PySceneDetect
Python and OpenCV-based scene cut/transition detection program & library.
DOVER
Video Quality Assessment on User Generated Contents from Aesthetic and Technical Perspectives.
ArtiMuse
Fine-Grained Image Aesthetics Assessment with Joint Scoring and Expert-Level Understanding.
图片 Image
Qwen-Image-2.1:
a unified text-to-image generation and image editing model in the Qwen family
GLM-Image:
an image generation model
OneTrainer:
One-stop solution for all your Diffusion training needs. Supports FLUX, Stable Diffusion 1.5/2.x/3.x/SDXL, Würstchen, PixArt, Hunyuan Video and more. Features full fine-tuning, LoRA, embeddings, masked training, automatic backups, and TensorBoard integration. GPL-3.0 licensed.
搜索 Search
OpenSearch GPT
SearchGPT / Perplexity clone, but personalised for you.
MindSearch
An LLM-based Multi-agent Framework of Web Search Engine (like Perplexity.ai Pro and SearchGPT).
nanoPerplexityAI
The simplest open-source implementation of perplexity.ai.
curiosity
Try to build a Perplexity-like user experience.
MiniPerplx
A minimalistic AI-powered search engine that helps you find information on the internet.
语音 Speech
kokoro:
https://hf.co/hexgrad/Kokoro-82M
Higgs Audio V2:
【Training】
KittenTTS:
Kitten TTS is an open-source realistic text-to-speech model with just 15 million parameters, designed for lightweight deployment and high-quality voice synthesis.
VibeVoice:
VibeVoice is a novel framework designed for generating expressive, long-form, multi-speaker conversational audio, such as podcasts, from text. It addresses significant challenges in traditional Text-to-Speech (TTS) systems, particularly in scalability, speaker consistency, and natural turn-taking.
FireRedTTS2:
FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot.
VoxCPM:
Open-sourced tokenizer-free multilingual speech synthesis model with high-quality TTS and style transfer workflows.
OmniVoice:
High-Quality Voice Cloning TTS for 600+ Languages
Whisper:
Whisper is a general-purpose speech recognition model that can be run locally offline. It can transcribe audio from and to multiple languages.
Step-Audio2:
Step-Audio 2 is an end-to-end multi-modal large language model designed for industry-strength audio understanding and speech conversation.
FunASR:
industrial-grade ASR toolkit; 170× realtime on GPU, 50+ languages, built-in VAD, punctuation, speaker diarization, and emotion detection. Includes non-autoregressive SenseVoice and LLM-based Fun-ASR-Nano models.
SenseVoice:
SenseVoice is a speech foundation model with multiple speech understanding capabilities, including automatic speech recognition (ASR), spoken language identification (LID), speech emotion recognition (SER), and audio event detection (AED).
世界模型 World Models
MIRA:
Multiplayer
Cosmos-3:
Cosmos is a world model development platform that consists of world foundation models, tokenizers and video processing pipeline to accelerate the development of Physical AI at Robotics & AV labs.
nano-world-model:
Minimalist, batteries-included repository for training video world models with diffusion-forcing, supporting long-horizon rollouts, 3D point-cloud generation, and model-predictive control with pretrained checkpoints (Simchowitz Lab, 700+ stars, MIT License, 2026)
https://github.com/shengshu-ai/minWM
world model
OpenWorldLib:
A unified codebase for world models, providing a standardized pipeline interface over existing open-source models (Matrix-Game-2, Hunyuan-GameCraft, FlashWorld, Cosmos-Predict-2.5, and others) across video generation, 3D scene generation, and reasoning. Apache-2.0.
龙虾 OpenClaw
NEXU:
The simplest desktop client for OpenClaw 🦞 — bridge your Agent to WeChat, Feishu, Slack & Discord in one click. Works with Claude Code, Codex & any LLM. BYOK, Oauth, local-first, chat from your phone 24/7.
统一模型 Unified Model
书籍 Book
《Build a Large Language Model (From Scratch)》
Implementing a ChatGPT-like LLM from scratch, step by step
《Understanding Deep Learning》
website with the book draft and Google Colabs of the book by Simon J.D. Prince
《Hands-On Large Language Models》
Covers LLM fundamentals, prompt engineering, and fine-tuning.
Foundations of Large Language Models
by Tong Xiao and Jingbo Zhu
《从零开始构建智能体》——从零开始的智能体原理与实践教程
📚 《从零开始构建智能体》——从零开始的智能体原理与实践教程
课程 Course
斯坦福 CS224N: Natural Language Processing with Deep Learning
This course provides a comprehensive insight into Deep Learning for NLP using PyTorch, emphasizing end-to-end neural models, eliminating the need for task-specific feature engineering, and equipping students with the skills to craft their own neural network solutions.
吴恩达: LLM series of courses
Focused courses on current generative AI engineering techniques.
llm-course: Course to get into Large Language Models (LLMs) with roadmaps and Colab notebooks.
Free LLM course with roadmap-style progression and practical Colab notebooks.
微软: Generative AI for Beginners
21 lessons covering generative AI fundamentals, prompt engineering, RAG applications, fine-tuning, and LLM app deployment with practical exercises.
斯坦福 CS25: Transformers United V4
This course delves into the transformative role of Transformers in deep learning, particularly their impact on the advancement of language models like ChatGPT and GPT-4.
普林斯顿 COS 597G (Fall 2022): Understanding Large Language Models
An advanced exploration into the transformative realm of LLMs, discussing state-of-the-art models, their profound capabilities, and associated challenges, with an emphasis on in-depth research, ethical considerations, and hands-on project experience, tailored for seasoned students versed in…
openai-cookbook
Examples and guides for using the OpenAI API.
Hands on llms
Learn about LLM, LLMOps, and vector DBS for free by designing, training, and deploying a real-time financial advisor LLM system.
LangGPT
Empowering everyone to become a prompt expert!
build nanoGPT
Video+code lecture on building nanoGPT from scratch.
LLM101n
Let's build a Storyteller.
Andrej Karpathy - Neural Networks: Zero to Hero
Build neural networks and language models from first principles.
Interactive visualization of Transformer
Interactive visualization of how transformer-based LLMs work, running a live GPT-2 model in the browser. #opensource
Anthropics:Prompt Engineering Interactive Tutorial
Prompt Engineering Interactive Tutorial by Anthropic
Cohere LLM University
free course on LLMs, embeddings, semantic search, and NLP applications.
Smol Vision
Recipes for shrinking, optimizing, customizing cutting edge vision models.
RAG++ : From POC to production
Advanced RAG course.
Weights & Biases AI Academy
Finetuning, building with LLMs, Structured outputs and more LLM courses.
LLM Evaluation: A Complete Course
Learn to build modern software with LLMs using the newest tools and techniques in the field.
HuggingFace Learn
Free hands-on courses using only open models.
RAG Techniques
This repository showcases various advanced techniques for Retrieval-Augmented Generation (RAG) systems. RAG systems combine information retrieval with generative models to provide accurate and contextually rich responses.
教程 Tutorial
Prompt Engineering Guide
This guide introduces Prompt Engineering, a discipline that optimizes interactions with Large Language Models, offering extensive resources, research, and tools.
Chip Huyen
ML Engineering, MLOps, and the use of ML in startups
Implementation of all RAG techniques in a simpler way.
Implementation of all RAG techniques in a simpler way
鱼皮的 Vibe Coding 零基础教程
程序员鱼皮的 AI 资源大全 + Vibe Coding 零基础教程,分享 OpenClaw 保姆级教程、大模型玩法(DeepSeek / GPT / Gemini / Claude / GLM)、最新 AI 资讯、Prompt 提示词大全、AI 知识百科(Agent Skills / RAG / MCP / A2A)、AI 编程教程(Harness Engineering)、AI 工具用法(Cursor / Claude Code / TRAE / Codex / Copilot)、AI 开发框架教程(Spring AI / LangChain)、AI 产品变现指南,帮你快速掌握 AI…
论文 Paper
The Llama 3 Herd of Models
(Meta, 2024-2025) - widely adopted open-weight family; default base for fine-tuning across NLP tasks.
Jamba: A Hybrid Transformer-Mamba Language Model
(2024) - hybrid Mamba-Transformer-MoE architecture.
Textbooks Are All You Need
Preprint
TÜLU 3: Pushing Frontiers in Open Language Model Post-Training
(AI2, 2024) - fully open post-training recipe with state-of-the-art results among open models.
Phi-4 Technical Report
(Microsoft, 2024) - small models trained on curated data, competitive with much larger ones on NLP benchmarks.
2 OLMo 2 Furious
(AI2, 2025) - fully open: weights, training data, code; reproducibility benchmark.
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
DeepSeek-R1-Zero, a model trained via large-scale reinforcement learning (RL) without supervised fine-tuning (SFT) as a preliminary step, demonstrated remarkable performance on reasoning.
Gemma 3 Technical Report
(Google, 2025) - 1B-27B open models with high local-to-global attention ratio to keep KV-cache tractable at 128K context.
Qwen3 Technical Report
Qwen3 is the large language model series developed by Qwen team, Alibaba Cloud.
Kimi K2 Technical Report
Kimi K2 is a state-of-the-art mixture-of-experts (MoE) language model with 32 billion activated parameters and 1 trillion total parameters.
LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
85M-Midtraining Data 22M Instruct Data
Olmo3
Charting a path through the model flow to lead open-source AI. Website
社区 Community
HuggingFace
Popular open platform for sharing ML models, datasets, and collaborating on NLP and generative AI projects.
模型上下文协议 MCP
mcp.so
Platform for MCP server resources and community.
modelcontextprotocol/servers
Anthropic's official reference MCP server implementations (GitHub, Slack, Postgres, Puppeteer, etc.). The authoritative source for understanding correct MCP server structure before building your own.
awesome-mcp-servers
Curated community list of MCP servers.
mcp.composio.dev
Connect Cursor, Windsurf, and Claude to 100+ fully managed MCP Servers with built-in auth; These servers are built by the community and are hosted by Composio
FastMCP
The fast, Pythonic way to build Model Context Protocol servers 🚀
FastAPI-MCP
Expose your FastAPI endpoints as Model Context Protocol (MCP) tools, with Auth!
技能 Skills
awesome-claude-skills
Curated list of Claude Skills, plugins, resources, and custom commands to extend terminal and API workflows. Apache 2.0 licensed.
Anthropics Skills
by Anthropic - Anthropic's official repository for Agent Skills — the SKILL.md format, a skill template, and example skills, the same format Claude Code loads natively.
awesome-claude-skills
security-flavoured skill list with design-engineering crossover
Skills.Sh
Open ecosystem by Vercel for installing reusable AI agent skills with a single command across 18+ platforms.
awesome-agent-skills
1000+ skills incl. design-md, enhance-prompt, react-components, shadcn-ui
claude-scientific-skills:
by K-Dense - "A set of ready-to-use Agent Skills for research, science, engineering, analysis, finance and writing." That's their description - modest, simple. That's how you can tell this is really one of the best skills repos on GitHub. If you've ever thought about getting a PhD... just read all…
LabClaw
Skill operating layer for biomedical AI agents with 211 production-ready SKILL.md files across 7 domains (biology, pharmacology, medicine, data science, literature search), enabling modular dry-lab reasoning and protocol composition for Stanford LabOS-compatible agents
推理 Open o1
https://github.com/marlaman/show-me
A visual and transparent alternative to open-source ChatGPT O1
g1
Using Llama-3.1 70b on Groq to create o1-like reasoning chains.
https://github.com/huggingface/open-r1
Fully open reproduction of DeepSeek-R1
https://github.com/Jiayi-Pan/TinyZero
Minimal reproduction of DeepSeek R1-Zero
LLaVA-o1:
code model
Marco-o1:
[code] [model]
https://github.com/simplescaling/s1
s1: Simple test-time scaling.
https://github.com/hiyouga/EasyR1
Efficient, scalable, multi-modality RL training framework based on veRL. Extends veRL to support vision-language models with GRPO algorithm for efficient RL training. Apache 2.0 licensed.
https://github.com/facebookresearch/swe-rl
Meta/UIUC/CMU
https://github.com/Liuziyu77/Visual-RFT
Shanghai AI Lab / SJTU
https://github.com/PeterGriffinJin/Search-R1
UIUC/Google
https://github.com/OpenManus/OpenManus-RL
UIUC/MetaGPT
AReaL:
AntGroup/Tsinghua
https://github.com/SkyworkAI/Skywork-OR1
Skywork AI
推理 Open o3
小语言模型 Small Language Model
https://github.com/jingyaogong/minimind
Train a 64M-parameter LLM from scratch in just 2 hours for $3. Complete from-scratch implementation covering MoE, data cleaning, pretraining, SFT, LoRA, RLHF (DPO/PPO/GRPO), tool use, and model distillation. All core algorithms implemented in pure PyTorch without high-level abstractions.…
https://github.com/loubnabnl/nanotron-smol-cluster
(使用Cosmopedia训练cosmo-1b)
https://github.com/allenai/OLMo
Open Language Model
https://github.com/skyzh/tiny-llm
learn LLM inference system on Apple Silicon for systems engineers: build a tiny vLLM + Qwen
https://huggingface.co/Nanbeige/Nanbeige4-3B-Thinking-2511
23T tokens预训练模型
小多模态模型 Small Vision Language Model
https://github.com/jingyaogong/minimind-v
🚀 「大模型」3小时从0训练27M参数的视觉多模态VLM!🌏 Train a 27M-parameter VLM from scratch in just 3 hours!
TinyLLaVA Factory
A Framework of Small-scale Large Multimodal Models.
Smol Vision
Recipes for shrinking, optimizing, customizing cutting edge vision models.
https://github.com/GeeeekExplorer/nano-vllm
Minimalist vLLM implementation in ~1,200 lines of Python. Educational yet performant with prefix caching, tensor parallelism, and CUDA graph acceleration. Comparable inference speeds to full vLLM. MIT licensed.
技巧 Tips
finetune-Qwen2-VL
Quick Start for Fine-tuning or continue pre-train Qwen2-VL Model.
https://github.com/jingyaogong/minimind
Train a 64M-parameter LLM from scratch in just 2 hours for $3. Complete from-scratch implementation covering MoE, data cleaning, pretraining, SFT, LoRA, RLHF (DPO/PPO/GRPO), tool use, and model distillation. All core algorithms implemented in pure PyTorch without high-level abstractions.…
LLM-Travel
致力于深入理解、探讨以及实现与大模型相关的各种技术、原理和应用
pytorch-llama
LLaMA 2 implemented from scratch in PyTorch.
Preference Optimization for Vision Language Models with TRL
【support model】
Distributed Training Guide
Best practices & guides on how to write distributed pytorch training code.
Related lists in Computer Science
See categoryTable of Contents
hesreallyhim/awesome-claude-code
A hand-picked collection of the finest of resources for the most awesome of agents, Claude Code, the undisputed champion of coding companions, from the unstoppable team…
Awesome Agent Skills
VoltAgent/awesome-agent-skills
A curated collection of 1000+ agent skills from official dev teams and the community, compatible with Claude Code, Codex, Gemini CLI, Cursor, and more.
Awesome Machine Learning
josephmisiti/awesome-machine-learning
A curated list of awesome Machine Learning frameworks, libraries and software.
Awesome Production Machine Learning
EthicalML/awesome-production-machine-learning
A curated list of awesome open source libraries to deploy, monitor, version and scale your machine learning
AWESOME DATA SCIENCE
academic/awesome-datascience
:memo: An awesome Data Science repository to learn and apply for real world problems.
Static Analysis
analysis-tools-dev/static-analysis
⚙️ A curated list of static analysis (SAST) tools and linters for all programming languages, config files, build tools, and more. The focus is on tools which improve…