Awesome AI Papers
A curated list of the most impressive AI papers
This page lists names, links and short descriptions. The original list on GitHub is the source and belongs to its authors.
2023 Papers >Computer Vision
Segment Anything (SAM)
Alexander Kirillov
2023 Papers >NLP
Toolformer: Language Models Can Teach Themselves to Use Tools (Toolformer)
by Meta AI, 2023 - A smaller model trained to translate human intention into actions (i.e. decide which APIs to call, when to call them, what arguments to pass, and how to best incorporate the results into future token prediction).
Sparks of Artificial General Intelligence: Early experiments with GPT-4 (GPT-4 Eval)
by Microsoft Research, 2023 - There are completely mind-blowing examples in the paper.
HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in HuggingFace (HuggingGPT)
Solving AI Tasks with ChatGPT and its Friends in HuggingFace
Generative Agents: Interactive Simulacra of Human (Gen Agents)
a paper that presents computational software agents that simulate believable human behavior
Tree of Thoughts: Deliberate Problem Solving with Large Language Models (ToT)
search over reasoning trees.
LIMA: Less Is More for Alignment (LIMA)
"less is more for alignment"; small high-quality SFT data goes a long way.
Voyager: An Open-Ended Embodied Agent with Large Language Models (Voyager)
Open-ended embodied agent in Minecraft
Mathematical discoveries from program search with large language models (FunSearch)
Nature, 2024. [All Versions]. Large language models (LLMs) have demonstrated tremendous capabilities in solving complex tasks, from quantitative reasoning to understanding natural language. However, LLMs sometimes suffer from confabulations (or hallucinations), which can result in them making…
2023 Papers >Audio Processing
2023 Papers >Multimodal Learning
AudioGPT: Understanding and Generating Speech, Music, Sound, and Talking Head (AudioGPT)
Understanding and Generating Speech, Music, Sound, and Talking Head [code] [demo]
ImageBind: One Embedding Space To Bind Them All (ImageBind)
CVPR'23, 2023. [All Versions]. [Project]. This work presents ImageBind, an approach to learn a joint embedding across six different modalities - images, text, audio, depth, thermal, and IMU data. The authors show that all combinations of paired data are not necessary to train such a joint…
2023 Papers >Reinforcement Learning
Mastering Diverse Domains through World Models (DreamerV3)
Danijar Hafner, Jurgis Pasukonis, Jimmy Ba, Timothy Lillicrap. Arxiv 2023; Key: DreamerV3, scaling property to world model; ExpEnv: deepmind control suite, atari, DMLab, minecraft
Efficient Online Reinforcement Learning with Offline Data (RLPD)
Philip J. Ball, Laura Smith, Ilya Kostrikov, and Sergey Levine. ICML, 2023.
Direct Preference Optimization: Your Language Model is Secretly a Reward Model (DPO)
Reframed preference alignment as a simple classification objective without explicit reward modelling.
2023 Papers >Other Papers
2022 Papers >Computer Vision
Patches Are All You Need (ConvMixer)
(Openreview'2021)
DreamFusion: Text-to-3D using 2D Diffusion (DreamFusion)
. 🌐 Project Page | 💻 Code
2022 Papers >NLP
Chain-of-Thought Prompting Elicits Reasoning in Large Language Models (CoT)
foundational result; intermediate reasoning steps improve performance.
Finetuned Language Models Are Zero-Shot Learners (FLAN)
finetuned language models as zero-shot learners.
Training language models to follow human instructions with human feedback (InstructGPT)
This paper presents an RLHF approach to using supervised learning to fine-tuning. It is also known as a paper that illustrates the kernel of ChatGPT's thinking. Presumably, ChatGPT is an extended version of InstructGPT that enables fine-tuning on larger datasets.
Training Compute-Optimal Large Language Models (Chinchilla)
by Hoffmann et al. at DeepMind. TLDR: introduces a new 70B LM called "Chinchilla" that outperforms much bigger LMs (GPT-3, Gopher). DeepMind has found the secret to cheaply scale large language models — to be compute-optimal, model size and training data must be scaled equally. It shows that most…
ReAct: Synergizing Reasoning and Acting in Language Models (ReAct)
The foundational paper defining the Thought/Action/Observation loop structure that underlies virtually every agent harness. Required reading for understanding why the loop is structured the way it is and where each harness component maps onto the reasoning-acting cycle.
BLOOM: A 176B-Parameter Open-Access Multilingual Language Model (BLOOM)
176B-parameter open multilingual LM, 46 natural languages.
Optimizing Language Models for Dialogue (ChatGPT)
Announcement of ChatGPT, a conversational model trained to answer follow-up questions, admit mistakes, challenge incorrect premises, and reject inappropriate requests. OpenAI blog, November 30, 2022.
2022 Papers >Audio Processing
2022 Papers >Multimodal Learning
2022 Papers >Reinforcement Learning
2022 Papers >Other Papers
Historical Papers
Neural Machine Translation by Jointly Learning to Align and Translate (RNNSearch-50)
This paper introduces an attention mechanism in RNNs to improve the long sequence modelling of RNNs. This paper introduces an attention mechanism to RNNs to improve their long sequence modelling capabilities. This enables RNNs to translate longer sentences more accurately.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding (BERT)
bidirectional transformer pretraining; foundation for most encoder-based NLP work since 2018. Read online with section navigation and the ACL source attached.
Highly accurate protein structure prediction with AlphaFold (Alphafold)
Nature, 2021. [All Versions]. This paper provides the first computational method that can regularly predict protein structures with atomic accuracy even in cases in which no similar structure is known. This approach is a canonical application of observation- and explanation- based method for…
Optimizing Language Models for Dialogue (ChatGPT)
Announcement of ChatGPT, a conversational model trained to answer follow-up questions, admit mistakes, challenge incorrect premises, and reject inappropriate requests. OpenAI blog, November 30, 2022.
Related lists in Computer Science
See categoryTable of Contents
hesreallyhim/awesome-claude-code
A hand-picked collection of the finest of resources for the most awesome of agents, Claude Code, the undisputed champion of coding companions, from the unstoppable team…
Awesome Agent Skills
VoltAgent/awesome-agent-skills
A curated collection of 1000+ agent skills from official dev teams and the community, compatible with Claude Code, Codex, Gemini CLI, Cursor, and more.
Awesome Machine Learning
josephmisiti/awesome-machine-learning
A curated list of awesome Machine Learning frameworks, libraries and software.
Awesome Production Machine Learning
EthicalML/awesome-production-machine-learning
A curated list of awesome open source libraries to deploy, monitor, version and scale your machine learning
AWESOME DATA SCIENCE
academic/awesome-datascience
:memo: An awesome Data Science repository to learn and apply for real world problems.
Static Analysis
analysis-tools-dev/static-analysis
⚙️ A curated list of static analysis (SAST) tools and linters for all programming languages, config files, build tools, and more. The focus is on tools which improve…