Skip to content
90

Awesome Prompts

Curated list of chatgpt prompts from the top-rated GPTs in the GPTs Store. Prompt Engineering, prompt attack & prompt protect. Advanced Prompt Engineering papers.

9k stars874 forks288 entriesLast push Sep 30, 2026 (today)License GPL-3.0

This page lists names, links and short descriptions. The original list on GitHub is the source and belongs to its authors.

Frameworks >Prompt Programming

DSPy

Write LM pipelines declaratively, then compile — DSPy auto-optimizes prompts and few-shot demonstrations. The strongest engineering-first approach.

In 11 listsDetails

Guidance

Interleave generation with constraints, regex/CFG, and control flow. Precision output control that goes beyond what prompts alone can achieve.

In 4 listsDetails

Frameworks >Automatic Prompt Optimization

TextGrad

Treats LLM feedback as "textual gradients" and backpropagates them to optimize prompts. Published in Nature.

In 2 lists

GEPA

Reflective Text Evolution — optimizes prompts, code, and agent configs. Claims +6–20 pts over GRPO on 6 tasks with fewer rollouts.

In 3 lists

Hermes Agent Self-Evolution

Evolutionary self-improvement for Hermes Agent — DSPy + GEPA (Genetic-Pareto Prompt Evolution) automatically evolves skills, tool descriptions, system prompts, and code via reflective search over execution traces (understands why things fail, not just that they failed); constraint gates (tests,…

In 2 lists

Reef

Open-source infrastructure for continually self-improving agents — connects agent inference, feedback, learning, and versioned delivery; either train model weights (Slime + SGLang) or improve the agent harness itself (prompts, rules, skills) from interaction feedback, no local GPUs required;…

Frameworks >Tool Use & Reliability

forge

Reliability layer for self-hosted LLM tool-calling — guardrails (rescue parsing, retry nudges, response validation), optional workflow constraints (required_steps, prerequisites, terminal_tool), and built-in eval suite. MIT, 2.2k+ stars, Feb 2026

reverify

Hallucination gate for agents — the model proposes claims, deterministic tools check each against ground truth and return VERIFIED/REFUTED with evidence; only what survives counts as fact. Ships as MCP server + CLI, with reverify rollover for lossless context handoff across resets. Caught every…

In 2 lists

Frameworks >Eval & Testing

promptfoo

Test-driven prompt engineering: regression tests, red teaming, model comparison, CI/CD integration. Acquired by OpenAI (Mar 2026) — remains open source.

In 11 listsDetails

OpenAI Evals

Open eval framework and benchmark registry — standardizes LLM performance measurement.

In 8 listsDetails

Terminal-Bench

Real-terminal agent benchmark (Stanford/Laude) — compile code, train models, set up servers in Docker-sandboxed environments; the de facto benchmark for agentic coding (2026).

In 2 lists

Frameworks >Red Team & Security

garak

LLM vulnerability scanner by NVIDIA — red teaming, prompt injection, jailbreak, and leakage detection.

In 4 listsDetails

OpenAI: Prompt Injection Defense

Official OpenAI guide on designing agents to resist prompt injection — browser agents, defense principles (2026).

The Promptware Kill Chain

Bruce Schneier (Harvard/Lawfare): reframes prompt injection as a 7-stage malware kill chain; 21/36 documented attacks already traverse 4+ stages. Featured at Black Hat 2026.

Microsoft Agent Governance Toolkit

7 packages (Python/Rust/TS/Go/.NET) — policy enforcement (<0.1ms), zero-trust agent identity (Ed25519 + SPIFFE), sandboxed execution; covers all OWASP Agentic Top 10; adapters for LangChain/CrewAI/ADK/OpenAI Agents SDK (Apr 2026)

In 5 listsDetails

agent-drift

Stress-test agents for goal drift and system-prompt violations across 6 value dimensions — multi-turn escalation, LLM-as-judge, interactive HTML reports; inspired by ICLR 2026 workshop paper (Apr 2026)

T3MP3ST

Autonomous red-teaming meta-harness for AI coding agents — recon → exploit → report against authorized targets, multi-agent offensive-security workflows, offline-model support; by elder-plinius (AGPL-3.0, 5.3k+ stars, July 2026)

In 3 lists

OpenAI Codex Security

Official OpenAI CLI and TypeScript SDK for finding, validating, and fixing security vulnerabilities — standard/deep scans, diff and working-tree targets, SARIF/CSV/JSON export, CI-native exit codes, pre-commit hooks (Apache-2.0, 8k+ stars, July 2026)

In 3 lists

SkillSpector

Security scanner for AI agent skills — detects vulnerabilities, malicious patterns, prompt injection, data exfiltration, and supply-chain risks in Claude Code, Codex, and MCP skills before installation (Apache-2.0, 14.7k+ stars, Mar 2026)

In 5 listsDetails

Frameworks >Eval & Observability

DeepEval

Unit testing for LLMs — G-Eval, hallucination, RAG faithfulness, agentic task metrics.

In 9 listsDetails

Langfuse

Open-source LLM engineering platform — tracing, evals, prompt management, A/B experiments.

In 10 listsDetails

Phoenix

Open-source AI observability & evaluation platform (Arize) — OpenTelemetry-native tracing for agents, LLM-as-judge evals, versioned datasets & experiments for prompt regression testing, prompt management with version control and replay, plus an MCP endpoint so Claude Code/Cursor can query traces…

In 10 listsDetails

Tracely-ai

Trace-native CI/CD for AI agents — grades every production trace as it lands (LLM-as-judge evaluators as trace-table columns), clusters failures into issues, freezes failing runs into hermetic replayable regression cases ($0 replay, no API keys), blocks the PR via CI gate, alerts via…

Frameworks >Low-Code & Workflow Platforms

Dify

Production-grade RAG and agent workflow platform — visual pipeline builder, multi-model support, plugin architecture.

In 14 listsDetails

Langflow

Drag-and-drop agent and chain builder — good for rapid prototyping of complex pipelines.

In 9 listsDetails

System Prompt Leaks

EliFuzz/awesome-system-prompts

Most comprehensive — Cursor, Devin, Windsurf, Claude Code, v0, Lovable, Perplexity, Manus, Replit, Warp and 20+ more. Actively maintained.

x1xhlol/system-prompts-and-models-of-ai-tools

20,000+ lines across 25+ tools (Claude Code, Cursor, Devin, Lovable, Manus, Windsurf, Kiro, v0, Codex, and more) — full tool definitions and internal agent logic; updated Mar 2026

In 3 lists

Piebald-AI/claude-code-system-prompts

Claude Code internal prompts — main system prompt, 18 tool descriptions, Plan/Explore/Task sub-agent prompts, 135+ version changelog

In 3 lists

asgeirtj/system_prompts_leaks

ChatGPT, Claude, Gemini system prompts and developer messages

In 4 listsDetails

jujumilk3/leaked-system-prompts

Well-organized, includes tool call constraints and persona definitions

elder-plinius/CL4R1T4S

Focused on Claude system prompt analysis

In 2 lists

Context Engineering

Effective Context Engineering for AI Agents — Anthropic

Anthropic's systematic guide to managing the full context state—system prompts, tools, MCP, and message history—as a finite, curated resource. Reframes harness design as "what configuration of context produces the desired behavior?" rather than just prompt wording.

In 5 listsDetails

Context Engineering Guide — Prompt Engineering Guide

davidkimai/Context-Engineering

first-principles handbook on context design, orchestration, and optimization

In 3 lists

Meirtz/Awesome-Context-Engineering

curated papers, frameworks, and implementation guides

In 2 listsDetails

muratcankoylan/Agent-Skills-for-Context-Engineering

comprehensive, MIT-licensed collection of Agent Skills for context engineering, multi-agent architectures, and production agent systems — context fundamentals/degradation/compression, memory systems, tool design, harness engineering, self-improvement loops; cited in academic research as…

In 4 listsDetails

NanoNets/Graft

Open-source codebase context layer for coding agents — builds a persistent, queryable code graph and pulls matching nodes into each prompt; 46% fewer tool calls, 42% token savings, 60% faster on a 162-run benchmark, 66% SWE-bench Verified (vs 54% cold); supports Claude Code, Cursor, Codex, Gemini…

In 2 lists

redhat-et/ripwire

"The ripgrep of AI context" (Red Hat): zero-dependency C++23 CLI + MCP server that lets coding agents find what they need without reading the whole repo — tree-sitter signatures at 74.7% fewer bytes than bodies, blast-radius queries, tests-to-run suggestions, and post-change verification with…

Agent Ecosystem >Frameworks

LangGraph

LangChain

CrewAI

CrewAI

Magentic-One

Microsoft

OpenAI Agents SDK

OpenAI

In 2 lists

OpenAI Agents SDK for JS/TS

OpenAI

In 2 lists

Claude Agent SDK

Anthropic

In 4 listsDetails

commerce-agents

Anthropic

GitHub Agentic Workflows (gh-aw)

GitHub

In 5 listsDetails

Google ADK

Google

In 4 listsDetails

Claude Code

Anthropic

In 2 lists

karpathy/autoresearch

Karpathy

In 5 listsDetails

Microsoft Agent Framework

Microsoft

In 4 listsDetails

openai/codex

OpenAI

In 9 listsDetails

DeerFlow 2.0

ByteDance

In 7 listsDetails

PilotDeck

OpenBMB / THUNLP / ModelBest / AI9Stars

AOS CE

Unicity

nanobot

HKUDS

In 6 listsDetails

OpenHuman

TinyHumans

In 3 lists

smolagents

HuggingFace

In 6 listsDetails

Flue

Astro

In 3 lists

Agno

Agno

In 7 listsDetails

browser-use

OSS

In 10 listsDetails

agent-browser

Vercel

In 2 lists

phone-harness

ShawnPana

Artemis

Google

Qwen-MM-Plugins

Alibaba/Qwen

In 2 lists

codebase-memory-mcp

DeusData

In 5 listsDetails

TencentDB Agent Memory

Tencent Cloud

In 3 lists

agentmemory

rohitg00

In 4 listsDetails

eve

Vercel

In 3 lists

Mastra

Gatsby team

In 10 listsDetails

PraisonAI

Mervin Praison

In 9 listsDetails

Portia AI

Portia Labs

Paperclip

Paperclip AI

In 3 lists

Goose

Block

In 3 lists

Gemini CLI

Google

In 9 listsDetails

kimi-code

Moonshot AI

In 3 lists

oh-my-codex

Yeachan Heo

In 3 lists

claw-code

UltraWorkers

Hermes Agent

Nous Research

In 7 listsDetails

herdr

herdr.dev

In 4 listsDetails

Orca

Stably

In 6 listsDetails

OpenSRE

Tracer Cloud

In 3 lists

DeepSeek Harness

DeepSeek

In 5 listsDetails

TrueForge

TrueFoundry

In 2 lists

OpenBot

CopilotKit

qm

YC Software

In 3 lists

Omnigent

Omnigent AI

In 5 listsDetails

Reef

Open-source infrastructure for continually self-improving agents — connects agent inference, feedback, learning, and versioned delivery; either train model weights (Slime + SGLang) or improve the agent harness itself (prompts, rules, skills) from interaction feedback, no local GPUs required;…

Dormice

BitMiracle AI

OpenConnector

Oomol

In 2 lists

Agent Ecosystem >Agent Skills

anthropics/skills

Official collection + spec (/spec/agent-skills-spec.md)

In 9 listsDetails

VoltAgent/awesome-agent-skills

1000+ community skills, works across all major platforms

In 3 listsDetails

vercel-labs/agent-skills

Vercel's official skills

In 2 lists

Agent Skills Docs — Anthropic

Official docs & spec

Equipping Agents for the Real World — Anthropic

Announcement post

In 3 lists

Skills vs MCP — LlamaIndex

When to use which

Agent Ecosystem >Harness Engineering

Harness Engineering — OpenAI

Official OpenAI post: "leveraging Codex in an agent-first world"

In 4 listsDetails

The Anatomy of an Agent Harness — LangChain

Component-by-component breakdown

In 2 lists

Improving Deep Agents with Harness Engineering — LangChain

TerminalBench 2.0 case study: 52.8% → 66.5%, same model

In 2 lists

The Importance of Agent Harness in 2026 — Philipp Schmid

"The harness is the dataset. Competitive advantage is the trajectories it captures."

Harness Engineering — Martin Fowler

Architecture perspective

In 2 lists

Skill Issue: Harness Engineering for Coding Agents — HumanLayer

Sub-agents as context firewalls, practical patterns

In 2 lists

Effective Harnesses for Long-Running Agents — Anthropic

Long-running agent design

In 4 listsDetails

SethGammon/Citadel

Production harness: 4-tier routing, parallel worktrees, lifecycle hooks, 6 skills

langchain-ai/deepagents

LangChain's opinionated deep agent harness (used in TerminalBench)

In 7 listsDetails

strukto-ai/mirage

Unified virtual filesystem for AI agents — mounts S3, GDrive, Slack, Gmail, Redis as one tree; agents use bash across every backend; Python/TypeScript SDKs, cache, snapshots (May 2026)

In 3 lists

Building a C Compiler with Parallel Claudes — Anthropic

How Anthropic used parallel Claude sub-agents to build a C compiler — generator/evaluator harness patterns

In 2 lists

QoderAI/better-harness

Open-source loop/harness improvement skill — turns project and session evidence into prioritized improvements and verifiable next steps for Claude Code, Codex, Cursor, and other coding agents (July 2026)

In 2 lists

harness-engineering

Ryan Lopopolo's anthology, field guide, and agent context bundle for harness engineering — shaping context and tools so agents can recover intent, operate systems, respect authority, prove outcomes, and leave the next run better equipped (CC-BY-4.0, July 2026)

In 2 lists

loop-engineering

"Stop prompting. Design the loop." — practical patterns, starters & CLI (loop-audit, loop-init, loop-cost) for systems that discover work, hand it to agents, verify results, and persist state across Claude Code, Codex, Grok, and OpenCode; report-only week one, scores loops on a "Loop Ready" rubric…

In 3 lists

ECC

The agent harness performance optimization system — skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond (MIT, 242k+ stars, Jan 2026)

In 7 listsDetails

SoL-Pi

NVIDIA's efficiency extension for the Pi coding agent — four opt-in harness mechanisms distilled from scaled auto-research loops (arXiv 2609.20519): Action Fusion (run follow-up validation in the same tool call), ObservationPack (stable handles with exact paged recall for repeated large outputs),…

In 2 lists

Papers >Foundations

Zero-Shot Reasoners (2022)

"Let's think step by step" — zero-shot CoT milestone

In 4 listsDetails

Self-Consistency (2022)

Multi-path sampling + majority vote: GSM8K 57% → 74%

In 5 listsDetails

ReAct (2023)

Reasoning + Acting interleaved — foundation of agent prompt design

In 7 listsDetails

APE: Human-Level Prompt Engineers (2023)

LLM auto-generates and selects instructions — beats human prompts

In 2 lists

A Prompt Engineering Universal Approximation Theorem (2026)

Formalizes prompt engineering as expressivity problem — proves a fixed Transformer backbone can approximate any continuous function by varying only the prompt; decomposes switching into routing/arithmetic/composition

Does Structured Intent Representation Generalize? A Cross-Language, Cross-Model Empirical Study of 5W3H Prompting (2026)

5W3H structured intent representation reduces cross-model output variance and avoids the dual-inflation bias of unstructured prompts; AI-expanded 5W3H matches manually crafted 5W3H across English, Japanese, and AI-assisted authoring

Papers >Automatic Optimization

ProTeGi / Gradient Descent for Prompts (2023)

Textual gradient descent — source paper for many auto-optimization methods

DSPy (2023)

Prompts as compilable programs — defines the engineering-first paradigm

MIPRO / Multi-Stage DSPy (2024)

Optimizes instructions and demonstrations across multi-stage LM programs

TextGrad (2024)

"Autograd for text" — LLM feedback as gradients, published in Nature

GEPA (2025)

Reflective evolution outperforms GRPO by 6–20 pts with fewer rollouts

Modular Prompt Optimization (2026)

Treats prompts as structured objects; optimizes each semantic section independently with local textual gradients

Causal Prompt Optimization (2026)

Reframes prompt design as causal estimation — uses Double Machine Learning to isolate prompt effects

Self-Evolving Memory for Prompt Optimization (2026)

Memory-augmented APO that stores historical refinement insights and reuses them across iterations

Combee: Scaling Prompt Learning for Self-Improving Agents (April 2026)

Berkeley/Stanford (Stoica, Zou, Gonzalez): scales parallel prompt learning with up to 17x speedup over ACE/GEPA via parallel scans and dynamic batching; evaluated on AppWorld, Terminal-Bench, FiNER

REprompt: Prompt Generation for Intelligent Software Development Guided by Requirements Engineering (Jan 2026)

Multi-agent prompt optimization framework that applies requirements engineering (elicitation, analysis, specification, validation) to generate production-ready system and user prompts for agent-based software development

Self-Distillation Improves Code Generation (April 2026)

Apple: embarrassingly simple self-distillation (SSD) — sample from model, fine-tune on raw unverified samples via cross-entropy; no reward model, no verifier, no RL; Qwen3-30B 42.4% → 55.3% pass@1 on LiveCodeBench v6; gains concentrate on hard problems; open source

SePO: Self-Evolving Prompt Agent for System Prompt Optimization (June 2026)

NUS/CityUHK: closes the self-referential loop by treating the prompt agent's own system prompt as an optimization target alongside task-agent prompts; open-ended evolutionary search with an archive of stepping-stone candidates; two-stage pre-train/fine-tune pipeline generalizes to held-out tasks;…

Papers >Reasoning Techniques

Chain of Draft (2025)

≤5 words per reasoning step — 91% of CoT accuracy at 7.6% of the tokens; 76% latency reduction

Thinking Without Words: Efficient Latent Reasoning with Abstract Chain-of-Thought (April 2026)

IBM Research AI: replaces verbal CoT with short sequences of learned, reserved vocabulary tokens; up to 11.6× fewer reasoning tokens with comparable accuracy on math, instruction-following, and multi-hop reasoning

Think Deep, Not Just Long (2026)

Longer CoT ≠ better reasoning — identifies "deep-thinking tokens" (high-revision tokens) as the true signal; enables cost-efficient test-time scaling

ReBalance: Efficient Reasoning with Balanced Thinking (2026)

Detects overthinking/underthinking via confidence variance and applies steering vectors to redirect reasoning — ICLR 2026; works on DeepSeek-R1, QwQ, o3-class models

InftyThink: Breaking Length Limits of Long-Context Reasoning (2026)

"Jagged" iterative reasoning — splits long reasoning into short segments with summaries, enabling unlimited depth without hitting context limits; ICLR 2026; +3–13% on MATH500/AIME24/GPQA

Reasoning Models Generate Societies of Thought (2026)

Google DeepMind: DeepSeek-R1/QwQ-32B superior reasoning emerges from simulating internal multi-agent dialogue — base models trained purely on reasoning accuracy spontaneously develop questioning, perspective-switching, and contradiction-resolving behaviors

Reasoning Theater: Disentangling Model Beliefs from CoT (2026)

For simple tasks, the model's final answer is already decodable from early-layer activations before CoT generates a single token — CoT produces genuine belief change only on hard problems; probe-guided early-exit reduces token generation by 80% on simple tasks

FLARE: Why Reasoning Fails to Plan (2026)

Diagnoses root cause of LLM agent long-horizon planning failures (stepwise reasoning induces greedy policy); FLARE (Future-aware Lookahead + Reward Estimation) lets LLaMA-8B surpass GPT-4o on planning benchmarks

Agentic Code Reasoning (March 2026)

Semi-formal reasoning using structured templates requiring explicit evidence — achieves 87% accuracy on code QA, 9 pp gain over standard agentic reasoning; enables interpretable code understanding for complex reasoning tasks

Reasoning Shift: How Context Silently Shortens LLM Reasoning (April 2026)

Contextual changes cause reasoning models to compress traces by up to 50%, reducing self-verification; simple problems unaffected but harder tasks suffer — critical finding for agent multi-turn reasoning

Rethinking Generalization in Reasoning SFT (April 2026)

Challenges "SFT memorizes, RL generalizes" — reasoning SFT with long CoT does generalize cross-domain, conditional on optimization dynamics; discovers safety-reasoning tradeoff (reasoning improves but safety degrades); 152 HF likes

RAGEN-2: Reasoning Collapse in Agentic RL (April 2026)

Identifies "template collapse" in agentic RL — models rely on fixed input-agnostic templates despite stable entropy; proposes mutual information (not entropy) as diagnostic for reasoning quality; Northwestern/Stanford/Microsoft; 49 HF likes

Optimality of LLMs on Planning Problems (April 2026)

Google DeepMind: first systematic study of whether LLMs produce optimal plans (not just valid); reasoning-enhanced LLMs significantly outperform classical satisficing planners (LAMA) in complex multi-goal configurations

Stratified Scaling Search for Test-Time in Diffusion Language Models (April 2026)

S³: inference-time procedure maintaining a population of partial denoising trajectories with verifier-based look-ahead and reward-tilted Gibbs distribution — first principled test-time scaling for discrete masked diffusion LMs

When to Think, When to Speak: Learning Disclosure Policies for LLM Reasoning (May 2026)

Side-by-Side (SxS) Interleaved Reasoning — makes disclosure timing a controllable decision in autoregressive generation; interleaves partial disclosures with continued private reasoning, releasing content only when supported by reasoning so far; improves accuracy–latency Pareto trade-offs on…

AI Co-Mathematician: Accelerating Mathematicians with Agentic AI (May 2026)

Google DeepMind: interactive workbench for open-ended mathematical research — ideation, literature search, computational exploration, theorem proving, theory building; manages uncertainty, tracks failed hypotheses, outputs native mathematical artifacts; scores 48% on FrontierMath Tier 4, a new…

Papers >Surveys

Survey of Automatic Prompt Engineering (2025)

Full overview of discrete / continuous / hybrid prompt optimization

Externalization in LLM Agents: Memory, Skills, Protocols, Harness (April 2026)

Comprehensive survey unifying memory, skills, protocols, and harness engineering as four forms of "cognitive externalization" — traces progression from weights → context → harness using cognitive artifact theory; Shanghai Jiao Tong / UCL

Beyond the Parameters: ICL to Causal RAG (April 2026)

Comprehensive survey treating context enrichment as a continuum — from in-context learning through RAG, GraphRAG, to CausalRAG; includes claim-audit framework and cross-paper evidence synthesis

Credit Assignment in Reinforcement Learning for Large Language Models (April 2026)

Comprehensive survey of credit assignment methods for LLM RL (reasoning + agentic) — covers 47 papers from Jan 2024 to Apr 2026; traces shift from reasoning-focused to agentic/multi-agent CA methods

Secure RAG: A Taxonomy of Attacks, Defenses, and Future Directions (April 2026)

Comprehensive taxonomy of RAG security — poisoning, extraction, membership inference, jailbreaks, and privacy leakage attacks with corresponding defense strategies and future research directions

Papers >RAG & Knowledge

GraphRAG (2025)

Graph-structured retrieval enabling multi-hop reasoning

Self-RAG (2024)

Model decides when and how to retrieve

In 2 lists

Agentic RAG Survey (2025)

Agents embedded in RAG pipelines — dynamic, reasoning-driven retrieval beyond static pipelines

A-RAG: Agentic RAG via Hierarchical Retrieval (2026)

Hierarchical retrieval interfaces enabling agents to dynamically navigate multi-level knowledge structures

In 2 lists

Procedural Knowledge at Scale Improves Reasoning (April 2026)

Meta AI: RAG for reasoning — decomposes trajectories into 32M reusable subquestion-subroutine pairs; retrieves procedural "how-to" knowledge within reasoning traces; +19.2% across math/science/coding

SoK: Agentic RAG — Taxonomy, Architectures, Evaluation (2026)

First Systematization of Knowledge for Agentic RAG — formalizes retrieval-generation loops as finite-horizon POMDPs; multi-dimensional taxonomy covering planning strategies, retrieval orchestration, memory paradigms, and tool coordination

LMM-Searcher: Long-horizon Agentic Multimodal Search (April 2026)

RUC: file-based visual context management + progressive on-demand image loading — scales to 100-turn search horizons, SOTA on MM-BrowseComp and MMSearch-Plus

Papers >Agent Reliability

Towards a Science of AI Agent Reliability (2026)

12 concrete reliability metrics across consistency, robustness, predictability, safety — capability gains ≠ reliability gains

In 2 lists

Agentic Reasoning for LLMs (2026)

Comprehensive survey: 3-layer framework (single-agent capabilities → self-evolving agents → multi-agent coordination); 202 Hugging Face likes

Why Do Web Agents Fail? A Hierarchical Planning Perspective (2026)

Decomposes web agent behavior into high-level planning, low-level grounding, and replanning — PDDL-structured plans outperform NL plans but grounding remains the dominant bottleneck; a single round of exploratory replanning substantially improves task success

Claw-Eval: Trustworthy Evaluation of Autonomous Agents (April 2026)

End-to-end evaluation suite with 300 human-verified tasks across 9 categories — trajectory-aware grading over 2,159 rubric items; finds vanilla LLM judges miss 44% of safety violations and 13% of robustness failures

TimeSeek: Temporal Reliability of Agentic Forecasters (April 2026)

Benchmark built from 150 regulated prediction markets evaluated at 5 lifecycle checkpoints — models are most competitive early and on high-uncertainty markets; search improves pooled accuracy but degrades 12% of conditions

ReliabilityBench: Evaluating LLM Agent Reliability Under Production-Like Stress (2026)

3D reliability surface R(k,ε,λ) unifying consistency, robustness, fault tolerance — chaos engineering for agents; ReAct outperforms Reflexion under stress; pass@1 overestimates reliability by 20–40%

Shepherd: A Runtime Substrate Empowering Meta-Agents with a Formalized Execution Trace (May 2026)

Stanford: Python substrate that makes agent execution a first-class object — typed events, Git-like trace, deterministic fork/replay/intervene primitives; 5× faster fork than Docker, >95% prompt-cache reuse; CooperBench pair-coding success 28.8% → 54.7%, 58% lower wall-clock on TerminalBench-2

EurekAgent: Agent Environment Engineering is All You Need For Autonomous Scientific Discovery (June 2026)

Tsinghua / Zhipu AI: argues the bottleneck in autonomous discovery is the environment, not the agent workflow — four environment-engineering dimensions (permissions, artifacts, budget, human-in-the-loop) enable off-the-shelf CLI agents to set SOTA on math, kernel engineering, and ML tasks at low…

AgentAtlas: Beyond Outcome Leaderboards for LLM Agents (May 2026)

UC Santa Cruz / MIT: six-state control-decision taxonomy and trajectory-failure vocabulary for separating outcome success from control-decision and trajectory quality; explicit label menus account for 14–40 pp of apparent agent capability

Stop Hand-Holding Your Coding Agent: Engineering the Loops that Replace Step-by-Step Prompting (July 2026)

Defines loop engineering as a new layer above prompt, context, and harness engineering — loop spec anatomy (trigger, goal, five-level verification ladder, architecture, stopping rule, memory), design principles, and anti-patterns from a corpus of 50 real loops

From Prompts to Contracts: Harness Engineering for Auditable Enterprise LLM Agents (July 2026)

Reconstructs prompt-dominant enterprise prototypes into code-owned, auditable harnesses — source-to-claim pipeline, seven validation dimensions (grounding, routing, trace, hygiene, recommendation language, runtime interfaces, latency), and the principle that "prompts are not guardrails"; validated…

Papers >Multi-Agent Coordination

Experience as a Compass: Multi-Agent RAG with Evolving Orchestration (April 2026)

HERA: 3-layer hierarchical framework that jointly evolves global orchestration strategies and local agent behaviors using experiential knowledge — role-aware prompt optimization drives targeted improvements for each agent's responsibilities

LangMARL: Natural Language Multi-Agent Reinforcement Learning (April 2026)

Brings credit assignment and policy gradient evolution from cooperative MARL into language space — enables LLM agents to autonomously evolve coordination strategies in dynamic environments

Agent Q-Mix: Selecting the Right Action for LLM Multi-Agent Systems (April 2026)

Reformulates topology selection as cooperative MARL — each agent selects communication actions that jointly induce round-wise communication graphs; improves coordination efficiency

Competition and Cooperation of LLM Agents in Games (April 2026)

LLM agents tend to cooperate in multi-round, non-zero-sum contexts rather than Nash equilibria — insights for designing cooperative multi-agent systems

G2CP: Graph-Grounded Communication Protocol for Multi-Agent Reasoning (2026)

Replaces free-text agent messages with explicit graph operations (traversal, subgraph fragments, updates) over a shared knowledge graph — 73% token reduction, 34% accuracy improvement, fully auditable reasoning chains

AdaptOrch: Task-Adaptive Multi-Agent Orchestration (2026)

Topology selection (parallel/sequential/hierarchical/hybrid) matters more than model choice — AdaptOrch automatically picks the right topology per task; 12–23% improvement over static single-topology baselines across SWE-bench, GPQA, and RAG

In 2 lists

The Orchestration of Multi-Agent Systems (2026)

Systematic academic analysis of MCP and A2A as complementary communication protocols; enterprise-grade multi-agent orchestration architecture covering governance, observability, and organizational adoption patterns

WebSwarm: Recursive Multi-Agent Orchestration for Deep-and-Wide Web Search (July 2026)

Progressive recursive delegation framework where search nodes pair local objectives with search modes, pass evidence upward, and recycle shared experience across sibling nodes; outperforms single-agent baselines on BrowseComp-Plus, WideSearch, DeepWideSearch, and GISA

Papers >Self-Improving Agents

Hyperagents: Self-Referential Meta-Agents (2026)

Meta FAIR: task agent and meta agent unified in a single editable program — meta layer can modify itself (recursive self-improvement); validated on code, paper review, robotics, and olympiad math; 2.1k HF likes; open source (facebookresearch/HyperAgents)

EvoSkills: Self-Evolving Agent Skills via Co-Evolutionary Verification (April 2026)

Skill Generator iteratively refines agent skills while a Surrogate Verifier co-evolves to provide actionable feedback without ground-truth; surpasses human-written skills on SkillsBench in 5 rounds; works on Claude Code and Codex

OpenClaw-RL: Train Any Agent Simply by Talking (2026)

Every agent interaction generates a next-state signal (user reply, tool output, GUI state) — OpenClaw-RL recovers all of them as live RL training sources via Hindsight-Guided On-Policy Distillation; one unified policy trains across conversation, terminal, SWE, and GUI tasks simultaneously (145 HF…

MetaClaw: Just Talk — An Agent That Meta-Learns and Evolves in the Wild (2026)

Continual meta-learning framework that jointly evolves a base LLM policy and a reusable skill library — skill-driven fast adaptation from failure trajectories + opportunistic gradient updates during idle periods; 21.4% → 40.6% accuracy on benchmarks (134 HF likes)

CORAL: Autonomous Multi-Agent Evolution for Open-Ended Discovery (April 2026)

Framework enabling autonomous multi-agent evolution via persistent memory, asynchronous execution, and collaborative exploration — 3–10x higher improvement rates with fewer evaluations than evolutionary baselines; 251 HF likes

SkillClaw: Collective Skill Evolution with Agentic Evolver (April 2026)

Cross-user trajectories continuously aggregated and refined by autonomous evolver into shared skill repository — collective skill evolution in multi-user agent ecosystems; 142 HF likes

SKILL0: In-Context Agentic RL for Skill Internalization (April 2026)

Progressively withdraws skill documentation during training until agents operate zero-shot — +9.7% on ALFWorld, +6.6% on Search-QA with <0.5k tokens per step; 133 HF likes

Memento-Skills: Let Agents Design Agents (2026)

Read-Write Reflective Learning over executable skill libraries — agents retrieve, execute, reflect, and rewrite their own skills without retraining the base model; evaluated on HLE and GAIA

Papers >Agent Safety

ClawSafety: "Safe" LLMs, Unsafe Agents (April 2026)

120 adversarial scenarios across 5 high-privilege domains (SWE/finance/medical/legal/DevOps), 3 injection channels (skill files, email, web); 40–75% attack success rate; safety depends on model + framework stack, not model alone

Supply-Chain Poisoning Attacks Against Agent Skill Ecosystems (April 2026)

DDIPE attack embeds malicious logic in skill documentation code examples; 1,070 adversarial skills across 15 MITRE ATT&CK categories; 11.6–33.5% bypass rate; responsible disclosure led to 4 confirmed vulnerabilities and 2 patches

BeSafe-Bench: Behavioral Safety Risks of Situated Agents (2026)

First benchmark across 4 real functional domains (Web, Mobile, Embodied VLM/VLA) with 9 safety-risk categories; even the best agent completes <40% of tasks under full safety constraints

Agents of Chaos (2026)

Two-week red-team study of live autonomous agents (email, Discord, shell, persistent memory) — documents 11 real attack categories including cross-agent unsafe practice propagation, identity spoofing, unauthorized resource consumption, and false task completion (32 HF likes)

LPS-Bench: Long-Horizon Safety Benchmarking for Computer-Use Agents (2026)

Safety benchmark for browser/computer-use agents focused on long-horizon tasks where risk accumulates across many UI actions — useful for testing confirmation discipline, phishing resistance, and context drift

Internal Safety Collapse in Frontier LLMs (2026)

Introduces TVD framework and ISC-Bench — frontier models fail at 95.3% rate on dual-use professional tasks where capability and harm co-occur; advanced models are more vulnerable than earlier LLMs because their capabilities become liabilities

Jailbreaking LLMs & VLMs: Mechanisms, Evaluation, and Unified Defense (2026)

First unified survey spanning both LLM and VLM jailbreak — covers template, in-context, RL, and multimodal attack types; proposes 3-layer defense framework (perception / generation / parameter layers)

Attack and Defense Landscape of Agentic AI (2026)

Dawn Song (UC Berkeley) et al. — first complete security survey for agentic AI systems (LLM + external tools/components); establishes threat model covering full attack surface and defense mechanisms; USENIX Security 2026

In 2 lists

Architecting Secure AI Agents: System-Level Defenses Against Indirect Prompt Injection (March 2026)

Greshake/Xiao/Suh et al. — security architecture paper arguing prompt injection must be handled at the system layer (permissioning, provenance, policy isolation), not by model alignment alone

Parallax: Why AI Agents That Think Must Never Act (April 2026)

Argues that prompt-based safety is architecturally insufficient for agents with execution capability; introduces Parallax, a plan-then-execute separation architecture with formal safety guarantees

Safety, Security, and Cognitive Risks in World Models (2026)

Comprehensive threat model for world-model-equipped agents — adversarial attacks, goal misgeneralisation, deceptive alignment, automation bias; extends MITRE ATLAS and OWASP to world model stack

Self-Propagating Attacks Across LLM Agent Ecosystems (March 2026)

Demonstrates how attacks can autonomously propagate across interconnected LLM agents — worm-like self-spreading malware targeting agent ecosystems via MCP, tool chains, and shared memory

From Untrusted Input to Trusted Memory: A Systematic Study of Memory Poisoning Attacks in LLM Agents (June 2026)

First systematic study of persistent memory poisoning — maps 4 write channels, 9 structural vulnerabilities, and 6 attack classes; introduces MPBench; shows current prompt-injection defenses are insufficient against cross-session memory manipulation

Agent Data Injection Attacks are Realistic Threats to AI Agents (July 2026)

New category of indirect prompt injection in which malicious data is disguised as trusted data (metadata, tool outputs, context structures, identifiers), bypassing existing IPI defenses; demonstrates real-world attacks on web and coding agents including Claude Code, Codex, and Gemini CLI

Papers >Medical & Health AI

Medical Reasoning with Large Language Models: A Systematic Review and Evaluation (April 2026)

Comprehensive review of medical reasoning methods + MR-Bench (real-world hospital data); reveals large gap between exam-level performance and authentic clinical decision-making

VeriSim: Evaluating Medical AI Under Realistic Patient Noise (April 2026)

Truth-preserving patient simulation framework injecting controllable, clinically evidence-grounded noise — evaluates medical AI robustness under realistic imperfect patient data conditions

Med-CAM: Minimal Evidence for Explaining Medical Decision Making (April 2026)

Minimal evidence extraction for medical AI explanations — identifies the smallest subset of input features sufficient for model decisions, improving interpretability without performance loss

ProMedical: Hierarchical Fine-Grained Criteria Modeling for Medical LLM Alignment (April 2026)

Hierarchical fine-grained criteria modeling for medical LLM alignment — structured clinical evaluation rubrics with multi-level criteria decomposition for improved medical reasoning and safety

Can Large Language Models Self-Correct in Medical Question Answering? (April 2026)

Exploratory study of LLM self-correction in medical QA — finds reflection can both correct and introduce errors; analyzes error correction dynamics across multiple reflection steps on MedQA, HeadQA, PubMedQA

Multi-Agent LLM Systems for Clinical Diagnosis: The Impact of Vendor Diversity (2026)

MIT/Harvard: mixed-vendor multi-agent diagnosis outperforms single-vendor teams — complementary inductive biases surface correct diagnoses that homogeneous teams miss; SOTA on RareBench and DiagnosisArena

Papers >Context & Memory

Active Context Compression (2026)

Focus agent architecture — autonomously consolidates history into a Knowledge block and prunes stale context; 22.7% token reduction on SWE-bench Lite, no accuracy loss

In 2 lists

Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models (2026)

ACE treats contexts as evolving playbooks with Generator/Reflector/Curator roles and incremental delta updates; defeats brevity bias and context collapse; +10.6% on agent benchmarks, +8.6% on finance; Stanford/CMU/Salesforce

Context Engineering: From Prompts to Corporate Multi-Agent Architecture (2026)

Defines context engineering as a standalone discipline for agentic AI; proposes a four-level maturity pyramid (Prompt Engineering → Context Engineering → Intent Engineering → Specification Engineering) and five context-quality criteria (relevance, sufficiency, isolation, economy, provenance)

AgeMem: Unified Long- and Short-Term Memory for LLM Agents (2026)

First to unify LTM (add/update/delete) and STM (retrieve/summarize/filter) as tool-based actions via GRPO RL; 7B model achieves +49.59% over no-memory baseline across 5 benchmarks; ICLR 2026 MemAgents Workshop

MSA: Memory Sparse Attention to 100M Tokens (2026)

End-to-end trainable sparse attention with linear complexity — scales to 100M tokens on 2×A800 GPUs with <9% degradation vs 16K baseline; Memory Interleaving enables multi-hop reasoning across scattered segments

Memory in the LLM Era: Modular Architectures in a Unified Framework (April 2026)

Decomposes agent memory into 4 modules (extraction, management, storage, retrieval); systematic benchmark comparison of all methods; composite design from existing modules surpasses prior SOTA

Are We Ready For An Agent-Native Memory System? (June 2026)

Tsinghua / HKUST / SJTU: first data-management study of agent memory — 12 systems + 2 baselines across 5 workloads and 11 datasets; four-module framework (representation/storage, extraction, retrieval/routing, maintenance); finds no single architecture dominates and localized maintenance…

ContextBench: A Benchmark for Context Retrieval in Coding Agents (2026)

First benchmark focused on whether coding agents retrieve the right repository context before editing — measures relevance, latency, and downstream task success under realistic codebase navigation pressure

Prompt Compression in the Wild (April 2026)

First large-scale empirical study of prompt compression trade-offs in production — 30K queries across multiple LLMs and 3 GPU classes; LLMLingua achieves up to 18% end-to-end speedup when prompt/ratio/hardware match; ECIR 2026; includes open-source profiler for latency break-even prediction

Thought-Retriever: Don't Just Retrieve Raw Data, Retrieve Thoughts for Memory-Augmented Agentic Systems (April 2026)

Memory mechanism that retrieves compressed reasoning "thoughts" rather than raw context — enables more efficient and reasoning-aware memory for long-horizon agents

GAM: Hierarchical Graph-based Agentic Memory for LLM Agents (April 2026)

Hierarchical graph-structured memory with role-aware modulation and temporal/confidence weighting; training-free, evaluated across multiple model scales

LongSeeker: Elastic Context Orchestration for Long-Horizon Search Agents (May 2026)

Context-ReAct paradigm with five atomic operations (Skip, Compress, Rollback, Snippet, Delete) for adaptive context management; proves expressive completeness of Compress; LongSeeker achieves 61.5% on BrowseComp and 62.5% on BrowseComp-ZH, substantially outperforming Tongyi DeepResearch and…

LLM Agents Are Latent Context Managers: Eliciting Self-Managed Context via a Proprioceptive Dashboard (July 2026)

VISTA: typed, addressable context blocks + runtime proprioceptive dashboard (token usage, recency, access history, context pressure) + recoverable full-fidelity archive; training-free and model-agnostic; raises Gemini-3-Flash from 22.7% to 50.7% on LOCA-Bench, with gains on BrowseComp-Plus and GAIA

Papers >Tool Use

CCTU: Tool Use under Complex Constraints (2026)

200-task benchmark across 12 constraint categories (resource, behavior, toolset, response) with step-level validation; no model exceeds 20% completion; models violate constraints in >50% of cases with limited self-correction

Agentic Tool Use in Large Language Models (April 2026)

Comprehensive framework for understanding tool use in agentic systems — schema understanding, calling conventions, error handling, tool composition patterns

Open, Reliable, and Collective: A Community-Driven Framework (April 2026)

OpenTools: standardized tool schemas and lightweight wrappers for plug-and-play use across agent frameworks; intrinsic evaluation suite tracking correctness, robustness, regressions

Act Wisely: Meta-Cognitive Tool Use in Agentic Multimodal Models (April 2026)

Alibaba: addresses meta-cognitive deficit where agents blindly invoke tools — HDPO framework reduces unnecessary tool invocations from 98% to 2% while increasing reasoning accuracy; first paper on "when NOT to use tools"

The Evolution of Tool Use in LLM Agents (2026)

Unified survey from single-tool call to multi-tool orchestration — covers reasoning-time planning, training/trajectory construction, safety, resource efficiency, open-environment completeness, and benchmark design (HIT & Harvard)

MCP-Atlas: Benchmarking LLM Agents on Real MCP Servers (2026)

Evaluates whether agents can use actual Model Context Protocol servers rather than toy tool schemas — measures correctness, protocol handling, and real-world MCP interoperability

Papers >Agent Evaluation

Signals: Trajectory Sampling and Triage for Agentic Interactions (April 2026)

Lightweight signal-based taxonomy for sampling informative agent trajectories post-deployment — 82% informativeness vs 54% random; organizes signals across interaction, execution, and environment dimensions; 6.2k HF likes

Agent Psychometrics: Task-Level Performance Prediction (April 2026)

Shifts evaluation from simple QA to multi-turn agentic assessment; newer benchmarks like SWE-bench Verified and Terminal-Bench test iterative agent behavior with execution feedback

YC-Bench: Benchmarking AI Agents for Long-Term Planning (April 2026)

Evaluates whether LLM agents maintain strategic coherence over long horizons — simulated startup over one-year horizon spanning hundreds of turns; tests consistent execution

When Users Change Their Mind: Evaluating Interruptible Agents (April 2026)

Tests agent ability to handle user interruptions during mid-task execution — critical requirement for realistic deployment in dynamic environments

SWE-CI: Evaluating Agents on Codebase Maintenance via CI (2026)

First CI-loop benchmark for long-term codebase maintainability — 100 tasks spanning 233 days and 71+ consecutive commits; shifts evaluation from static single-fix to dynamic long-horizon reasoning

SWE-Skills-Bench (2026)

565 real-world SE tasks measuring whether agent skills actually improve outcomes — 39/49 public skills give zero gain; average improvement only +1.2%; reveals fundamental gap in skill design

LongCLI-Bench: A Benchmark for Long-Horizon Agentic Programming in the CLI (2026)

Benchmarks terminal-based coding agents on long-horizon programming tasks that require sustained planning, repo navigation, debugging, and recovery over many steps instead of single-fix patches

ProjDevBench: Benchmarking AI Agents on End-to-End Software Project Development (2026)

Evaluates whether agents can build complete software projects from requirements to implementation and validation, rather than solving isolated bug-fix tasks; targets end-to-end project delivery realism

LiveClawBench: Benchmarking LLM Agents on Complex, Real-World Assistant Tasks (April 2026)

Evaluates agents on compositional, real-world assistant tasks requiring planning, tool use, and recovery — closer to production deployment scenarios than static QA benchmarks

RiskWebWorld: GUI Agents in E-commerce Risk Management (April 2026)

Realistic interactive benchmark for GUI agents in high-stakes professional workflows — 100 real-world e-commerce risk scenarios testing sequential decision-making under uncertainty

OccuBench: Real-World Professional Tasks via Language World Models (April 2026)

100 professional task scenarios across 10 industries and 65 domains — evaluates AI agents on realistic occupational workflows using language world models for environment simulation

In 2 lists

EpiBench: Multi-turn Research Workflows for Multimodal Agents (April 2026)

Benchmarks multimodal agents on episodic scientific research workflows — literature search, figure extraction, cross-paper synthesis; built on smolagents with persistent memory and tool use

Ask Early, Ask Late, Ask Right: When Does Clarification Timing Matter for Long-Horizon Agents (May 2026)

First forced-injection framework measuring how clarification value changes over the execution trajectory across goal/input/constraint/context dimensions; 6,000+ runs, 4 frontier models, 3 benchmarks; finds goal clarifications lose nearly all value after 10% execution, input clarifications retain…

Reasoning Is Not Free: Robust Adaptive Cost-Efficient Routing for LLM-as-a-Judge (May 2026)

ICML 2026: controlled comparisons show reasoning judges substantially improve accuracy on structured-verification tasks (math, coding) but yield limited or negative gains on simpler evaluations while costing significantly more compute; proposes RACER, a distributionally-robust routing policy that…

Papers >Instruction Following

MOSAIC: Granular Instruction Following Evaluation (2026)

Modular benchmark with up to 20 application-oriented generation constraints per prompt; finds compliance degrades with constraint count and position (primacy/recency bias) — exposes multi-instruction conflict effects

Rubrics to Tokens: Token-Level Rewards for Instruction Following (April 2026)

Rubric-based RL with Token-Level Relevance Discriminator — solves credit assignment for instruction following by predicting which tokens satisfy specific constraints; fine-grained optimization

Schema Key Wording as an Instruction Channel in Structured Generation (April 2026)

Discovers that schema key wording itself acts as an implicit instruction signal under constrained decoding — changing JSON key names alters model behavior even when semantic content is identical

One Token Away from Collapse: Fragility of Instruction-Tuned Helpfulness (April 2026)

Trivial lexical constraints (banning one punctuation mark) cause 14–48% response collapse in instruction-tuned LLMs — identified as planning failure via mechanistic analysis; base models show no collapse

Instruction Bleed: Cross-Module Interference in Prompt-Composed Agentic Systems (June 2026)

Formalizes Compositional Behavioral Leakage (CBL) — prompt modules sharing a context window silently shift each other's behavior; introduces a three-channel perturbation protocol (volume / content / form) and detects Cohen's d = 0.63 content-channel interference in a deployed job-evaluation agent;…

Enforcing Hierarchical Instruction-Following via Neuro-Symbolic Alignment (April 2026)

NSHA: formulates hierarchical instruction resolution as constraint satisfaction, solved with SAT solver-guided inference-time reasoning — resolves conflicts between system prompts, user instructions, and tool outputs

DEFT: Distribution-guided Efficient Fine-Tuning for Human Alignment (April 2026)

Distribution-guided efficient fine-tuning for alignment — uses data distribution properties to guide selective parameter updates, improving alignment quality with reduced compute

Papers >Multimodal Prompting

S-Agent: Spatial Tool-Use Elicits Reasoning for Spatial Intelligence (June 2026)

Spatial reasoning as spatio-temporal evidence accumulation — VLM planner + hierarchical 2D/3D spatial tools + dual memory; training-free gains on open-source and closed-source VLMs; S-Agent-8B matches GPT-5.4 and Gemini 3 on spatial benchmarks

Graph-of-Mark: Spatial Reasoning via Visual Prompting (2026)

Overlays scene graphs onto input images at the pixel level to model object relationships — up to +11 percentage points on VQA and localization across 4 datasets, zero-shot

Look Twice: Training-Free Evidence Highlighting in MLLMs (April 2026)

Inference-time framework exploiting MLLM attention patterns to identify relevant visual regions and text, then re-conditions generation on highlighted evidence — consistent VQA improvements, no training required

Agentic-MME: What Agentic Capability Really Brings to Multimodal Intelligence? (April 2026)

Systematic evaluation of agentic capability in multimodal LLMs — decomposes tasks into perception, reasoning, and action levels; reveals where agentic loops help vs. where they add overhead

FeynmanBench: Diagrammatic Physics Reasoning for MLLMs (April 2026)

First benchmark for Feynman diagram tasks — evaluates multistep diagrammatic reasoning requiring conservation laws, symmetry constraints, and graph topology; 2000+ tasks across Standard Model interactions

MERRIN: Multimodal Evidence Retrieval in Noisy Web Environments (April 2026)

Benchmark for multimodal evidence retrieval and multi-hop reasoning over noisy web content — even strongest agent (Gemini-3.1-Pro) achieves only 40.1%; finds more search ≠ better performance

Zooming without Zooming: Region-to-Image Distillation for Fine-Grained Multimodal Perception (2026)

Converts inference-time zooming into training-time primitive — teaches MLLMs fine-grained perception in single forward pass; introduces ZoomBench (845 VQA across 6 perceptual dimensions); SOTA on fine-grained benchmarks

Papers >Embodied AI & World Models

VLA-World: Vision-Language-Action World Models for Autonomous Driving (April 2026)

Unifies predictive imagination with reflective reasoning for driving foresight — action-derived trajectory guides next-frame generation, then reasons over the imagined frame to refine planning

EmbodiedClaw: Conversational Workflow Execution for Embodied AI Development (April 2026)

Conversational framework for embodied AI development — batch simulation environment synthesis, automatic scene creation, controllable scene editing, and workflow execution via natural language

StarVLA: Lego-like Codebase for VLA Model Development (April 2026)

Open-source modular VLA framework — swappable backbone (VLM/world-model) and action heads, cross-embodiment learning, unified evaluation across LIBERO, SimplerEnv, RoboTwin, RoboCasa, BEHAVIOR-1K

Human-to-Robot Imitation Learning: A Survey and Taxonomy of Methods (April 2026)

Comprehensive survey of human-to-robot imitation learning — behavioral cloning, inverse reinforcement learning, adversarial imitation, and their combinations; includes taxonomy, benchmarks, and open challenges

The Great March 100: 100 Detail-oriented Tasks for Evaluating Embodied AI Agents (2026)

100 detail-oriented embodied AI tasks spanning manipulation, navigation, and reasoning — evaluates fine-grained physical world understanding beyond coarse task completion

VLA-Forget: Vision-Language-Action Unlearning for Embodied Foundation Models (April 2026)

First unlearning method for VLA models — removes target behaviors while preserving general capabilities; introduces forget/retain/boundary splits and real-robot OXE benchmarks

Papers >Voice & Realtime Agents

Building Enterprise Realtime Voice Agents from Scratch (2026)

Salesforce AI Research: complete tutorial for production voice agents — cascaded streaming pipeline (STT→LLM→TTS), ~750ms TTFA, function calling, full open-source codebase with 9 chapters

VoiceMem: Streaming Dual-Brain Memory for Real-Time Interaction (Aug 2026)

Tsinghua: long-term memory for realtime voice agents — "left brain" stores compressed facts (Mem0-level accuracy at ~300 tokens/query), "right brain" tracks emotional attribution across nodes; fully streaming architecture with speculative prefetch keeps added latency near zero; open model family +…

Tools & Libraries

LangChain

LLM orchestration and chaining

In 20 listsDetails

LlamaIndex

Data ingestion and RAG pipelines

In 14 listsDetails

anydoc

Convert Word, PowerPoint, Excel, OpenDocument, RTF, EPUB, CSV, and PDF to clean Markdown — Rust core with Node.js/Python bindings; agent/RAG document ingestion (Aug 2026)

HyperFrames

Open-source HTML-to-video rendering framework built for agents — write HTML/CSS with seekable animations and render deterministic MP4s; agent skills, CLI, and hosted authoring workflows (Mar 2026)

In 2 lists

headroom

Compress tool outputs, logs, files, and RAG chunks before they reach the LLM — 20% fewer tokens for coding agents, 60–95% fewer for JSON; ships as a library, proxy, and MCP server (Jan 2026)

In 3 lists

OptMem

Permanent memory for AI agents — append-only log + binary-tree summaries, 426-token prompt, plug-and-play with Claude Code/Codex/etc. via AGENTS.md/CLAUDE.md (July 2026)

graphify

Turn any codebase — plus docs, SQL schemas, configs, and PDFs — into a queryable knowledge graph. Ships as a /graphify skill for Claude Code, Cursor, Codex, and Gemini CLI; local deterministic AST parsing, every edge explained, no vector store (Apr 2026)

In 4 listsDetails

LiteLLM

Unified API for 100+ LLM providers

In 16 listsDetails

Ollama

Run LLMs locally — desktop app, multimodal, structured outputs

In 12 listsDetails

Semantic Kernel

Microsoft's LLM SDK — now merging with AutoGen into Microsoft Agent Framework (2026)

In 11 listsDetails

TensorZero

LLM gateway + observability + optimization

In 3 lists

Outlines

Structured text generation and constrained outputs

In 6 listsDetails

PydanticAI

Official Pydantic agent runtime — typed tools, structured outputs, evals, production-ready (V1 stable)

In 12 listsDetails

Instructor

Most widely used library for structured LLM outputs — typed extraction from any model, 3M+ monthly downloads

In 2 lists

qwen-audio-agent

Realtime voice runtime for AI agents — keeps agents talking, working, and present while they think or use tools (no dead air during tool calls); pluggable STT/TTS and realtime providers, embeddable gateway, TUI/desktop apps, Agent Skills support, ACP-compatible; works with Claude Code, Codex,…

In 2 lists

LM Evaluation Harness

EleutherAI's unified LLM evaluation framework

In 7 listsDetails

Weights & Biases

Experiment tracking and LLMOps

Promptingguide.ai

Comprehensive prompt engineering reference (DAIR-AI)

In 5 listsDetails

awesome-ai-agents-2026

Most comprehensive list of 2026 AI agents, frameworks & tools — 300+ resources, 20+ categories, updated monthly

Awesome-Agent-Papers

Curated papers on LLM agents: methodology, applications, challenges — covers STRIDE, planning, tool use, memory, multi-agent (2026)

Awesome-Agentic-Reasoning

Papers and resources on agentic reasoning from foundational to multi-agent coordination — 3-layer framework (2026)

Agent-Memory-Paper-List

Curated papers on memory architectures for LLM agents — long-term, short-term, attention mechanisms (2026)

awesome-ai-agent-papers

Curated 2025–2026 papers on agent engineering, memory, eval, and workflows

In 3 listsDetails

langgptai/awesome-claude-prompts

Claude-optimized prompts — XML tags, extended thinking, long-context patterns

langgptai/awesome-deep-research-prompts

Prompts for OpenAI Deep Research, Gemini Deep Research, Perplexity Labs

ML-GSAI/Diffusion-LLM-Papers

Curated papers on diffusion language models — LLaDA, Dream, MMaDA, consistency sampling, fast inference; 169 stars, actively maintained (2026)

Anthropic Prompt Library

Official production-ready prompts from Anthropic

ai-engineering-from-scratch

The most complete open-source AI engineering curriculum — 523 lessons / 20 phases / ~342 hours; dedicated prompt engineering, agent engineering, MCP, and Agent Skills phases where every lesson ships a reusable artifact (prompt, skill, agent, MCP server); Python/TypeScript/Rust, MIT, 55k+ stars,…

In 3 lists

NirDiamant/Prompt_Engineering

22 Jupyter Notebook tutorials from basics to advanced — CoT, few-shot, templates, multi-language

In 5 listsDetails

automotive-skills-suite

152 installable Claude skills for automotive engineering — ISO 26262, ISO/SAE 21434, ISO 21448 SOTIF, AIAG-VDA, ASPICE, AUTOSAR; builder + reviewer pairs with xlsx deliverables

See category
94

Awesome OpenClaw Skills

VoltAgent/awesome-openclaw-skills

The awesome collection of OpenClaw skills. 5,400+ skills filtered and categorized from the official OpenClaw Skills Registry.🦞

Fresh★ 53k830 entriesPushed today
92

Awesome DeepSeek Harness (DSH) Plugin

awesome-dsh-plugin/awesome-dsh-plugin

A curated list of plugins for DeepSeek Harness (dsh) · DeepSeek Harness 插件精选列表

Fresh★ 17k1654 entriesPushed today
91

Awesome Guidelines

Kristories/awesome-guidelines

Programming style, best practices, and coding conventions.

Fresh★ 11k166 entriesPushed 2 days ago
90

Awesome

sindresorhus/awesome

😎 Awesome lists about all kinds of interesting topics [NOTE: Pull requests are temporarily disabled until I have a chance to catch up with the existing ones]

Fresh★ 513k51 entriesPushed 28 days ago
90

Awesome README

matiassingers/awesome-readme

A curated list of awesome READMEs

Fresh★ 22k143 entriesPushed yesterday
90

Awesome C

oz123/awesome-c

A curated list of awesome C frameworks, libraries, resources and other shiny things. Inspired by all the other awesome-... projects out there.

Fresh★ 12k557 entriesPushed 10 days ago