[code]
Awesome LLM Reasoning
From Chain-of-Thought prompting to OpenAI o1 and DeepSeek-R1 🍓
This page lists names, links and short descriptions. The original list on GitHub is the source and belongs to its authors.
Survey >2025
Survey >2024
Survey >2022
Analysis >2025
Analysis >2024
Analysis >2023
Large Language Models Can Be Easily Distracted by Irrelevant Context.
by Google Research et al., 2023 - Adding the instruction "Feel free to ignore irrelevant information given in the questions." consistently improves robustness to irrelevant context.
On Second Thought, Let's Not Think Step by Step! Bias and Toxicity in Zero-Shot Reasoning.
Bias and Toxicity in Zero-Shot Reasoning.
Analysis >2022
Analysis >2025
s1: Simple test-time scaling.
(2025) - small open reasoning recipe via budget-forcing.
Analysis >2024
PPM: Automated Generation of Diverse Programming Problems for Benchmarking Code Generation Models
[code] Simin Chen, XiaoNing Feng, Xiaohong Han, Cong Liu, Wei Yang FSE'24
OpenAI o1.
(2024-2025) - test-time-compute reasoning systems.
Let's Verify Step by Step.
process-supervised reward models for reasoning.
REFINER: Reasoning Feedback on Intermediate Representations.
[project] [code]
Analysis >2023
Synthetic Prompting: Generating Chain-of-Thought Demonstrations for Large Language Models.
Generating Chain-of-Thought Demonstrations for Large Language Models.
Rethinking with Retrieval: Faithful Large Language Model Inference.
by University of Pennsylvania et al., 2022 - They shows the potential of enhancing LLMs by retrieving relevant external knowledge based on decomposed reasoning steps obtained through chain-of-thought (CoT) prompting. I predict we're going to see many of these types of retrieval-enhanced LLMs in…
PAL: Program-aided Language Models.
[project] [code]
Self-consistency improves chain of thought reasoning in language models.
Multi-path sampling + majority vote: GSM8K 57% → 74%
Analysis >2022
Large Language Models are Zero-Shot Reasoners.
"Let's think step by step" — zero-shot CoT milestone
Analysis >2025
Introducing Visual Perception Token into Multimodal Large Language Model.
[code] [model] [dataset]
LlamaV-o1: Rethinking Step-by-step Visual Reasoning in LLMs.
[project] [code] [model]
Embodied-Reasoner: Synergizing Visual Search, Reasoning, and Action for Embodied Interactive Tasks
[project][code][dataset] Wenqi Zhang, Mengna Wang, Gangao Liu, Xu Huixin, Yiwei Jiang, Yongliang Shen, Guiyang Hou, Zhe Zheng, Hang Zhang, Xin Li, Weiming Lu, Peng Li, Yueting Zhuang. Preprint'25
Analysis >2024
Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models.
[project] [code]
Analysis >2023
MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action.
[project] [code] [demo]
ViperGPT: Visual Inference via Python Execution for Reasoning.
[project] [code]
Analysis >2025
Analysis >2024
Analysis >2023
Teaching Small Language Models to Reason.
They finetune a student model on the chain of thought (CoT) outputs generated by a larger teacher model. For example, the accuracy of T5 XXL on GSM8K improves from 8.11% to 21.99% when finetuned on PaLM-540B generated chains of thought.
Analysis >2022
Scaling Instruction-Finetuned Language Models.
by Google - They find that instruction finetuning with the above aspects dramatically improves performance on a variety of model classes (PaLM, T5, U-PaLM), prompting setups (zero-shot, few-shot, CoT), and evaluation benchmarks. Flan-PaLM 540B achieves SoTA performance on several benchmarks. They…
Other Useful Resources
Agent Shadow Brain
Self-evolving AI coding intelligence with infinite memory (TurboQuant), genetic algorithm self-evolution, predictive bug detection, PageRank knowledge graphs, swarm intelligence, and adversarial defense.
LLM Reasoners
A library for advanced large language model reasoning.
Chain-of-Thought Hub
Benchmarking LLM reasoning performance with chain-of-thought prompting.
Omni Skills Forge
50,000+ curated AI agent skills for Claude Code, Cursor, Copilot, Windsurf, Cline. Visual dashboard, one-click install, skill doctor, auto-update.
ThoughtSource
Central and open resource for data and tools related to chain-of-thought reasoning in large language models.
AgentChain
Chain together LLMs for reasoning & orchestrate multiple large models for accomplishing complex tasks.
google/Cascades
Python library which enables complex compositions of language models such as scratchpads, chain of thought, tool use, selection-inference, and more.
LogiTorch
PyTorch-based library for logical reasoning on natural language.
salesforce/LAVIS
One-stop Library for Language-Vision Intelligence.
facebookresearch/RAM
A framework to study AI models in Reasoning, Alignment, and use of Memory (RAM).
Related lists in Miscellaneous
See categoryAwesome OpenClaw Skills
VoltAgent/awesome-openclaw-skills
The awesome collection of OpenClaw skills. 5,400+ skills filtered and categorized from the official OpenClaw Skills Registry.🦞
Awesome DeepSeek Harness (DSH) Plugin
awesome-dsh-plugin/awesome-dsh-plugin
A curated list of plugins for DeepSeek Harness (dsh) · DeepSeek Harness 插件精选列表
Awesome Guidelines
Kristories/awesome-guidelines
Programming style, best practices, and coding conventions.
Awesome
sindresorhus/awesome
😎 Awesome lists about all kinds of interesting topics [NOTE: Pull requests are temporarily disabled until I have a chance to catch up with the existing ones]
Awesome Prompts
ai-boost/awesome-prompts
Curated list of chatgpt prompts from the top-rated GPTs in the GPTs Store. Prompt Engineering, prompt attack & prompt protect. Advanced Prompt Engineering papers.
Awesome README
matiassingers/awesome-readme
A curated list of awesome READMEs