Skip to content
83

Awesome AI Agent Papers

A curated collection of AI agent research papers released in 2026, covering agent engineering, memory, evaluation, workflows, and autonomous systems.

1.8k stars187 forks388 entriesLast push Sep 21, 2026 (9 days ago)License MIT

This page lists names, links and short descriptions. The original list on GitHub is the source and belongs to its authors.

Multi-Agent (54)

Agon: An Autonomous Large-Scale Omnidisciplinary Research System Built on Prompt Economy

PerceptUI: LLM Agents as Human-Aligned Synthetic Users for UI/UX Evaluation

OpenCLAW-P2P v6.0: Resilient Multi-Layer Persistence, Live Reference Verification, and Production-Scale Evaluation of…

AutoNumerics: An Autonomous, PDE-Agnostic Multi-Agent Pipeline for Scientific Computing

Beyond Offline A/B Testing: Context-Aware Agent Simulation for Recommender System Evaluation

CityReal: Human-Aligned Urban Behavior and City Dynamics Simulation with Large-Scale LLM Agents

CORAL: Towards Autonomous Multi-Agent Evolution for Open-Ended Discovery

DyTopo: Dynamic Topology Routing for Multi-Agent Reasoning via Semantic Matching

RuleSmith: Multi-Agent LLMs for Automated Game Balancing

CommCP: Efficient Multi-Agent Coordination via LLM-Based Communication with Conformal Prediction

AgenticPay: A Multi-Agent LLM Negotiation System for Buyer-Seller Transactions

Gender Dynamics and Homophily in a Social Network of LLM Agents

ROMA: Recursive Open Meta-Agent Framework for Long-Horizon Multi-Agent Systems

ORCH: many analyses, one merge — a deterministic multi-agent orchestrator

H-AdminSim: A Multi-Agent Simulator for Realistic Hospital Administrative Workflows

Agyn: A Multi-Agent System for Team-Based Autonomous Software Engineering

Multi-Agent Teams Hold Experts Back

Evolving Interpretable Constitutions for Multi-Agent Coordination

Scaling Multiagent Systems with Process Rewards

MonoScale: Scaling Multi-Agent System with Monotonic Improvement

Task-Aware LLM Council with Adaptive Decision Pathways for Decision Support

SYMPHONY: Synergistic Multi-agent Planning with Heterogeneous Language Model Assembly

Learning to Recommend Multi-Agent Subgraphs from Calling Trees

Learning Decentralized LLM Collaboration with Multi-Agent Actor Critic

AgenticSimLaw: A Juvenile Courtroom Multi-Agent Debate Simulation for Explainable High-Stakes Tabular Decision Making

Epistemic Context Learning: Building Trust the Right Way in LLM-Based Multi-Agent Systems

Adaptive Confidence Gating in Multi-Agent Collaboration for Efficient and Optimized Code Generation

CASTER: Context-Aware Strategy for Task Efficient Routing in Multi-Agent Systems

Phase Transition for Budgeted Multi-Agent Synergy

Dynamic Role Assignment for Multi-Agent Debate

Learning to Collaborate: An Orchestrated-Decentralized Framework for Peer-to-Peer LLM Federation

Mixture-of-Models: Unifying Heterogeneous Agents via N-Way Self-Evaluating Deliberation

Multi-Agent Constraint Factorization Reveals Latent Invariant Solution Structure

MAS-Orchestra: Understanding and Improving Multi-Agent Reasoning Through Holistic Orchestration and Controlled…

MASCOT: Towards Multi-Agent Socio-Collaborative Companion Systems

If You Want Coherence, Orchestrate a Team of Rivals: Multi-Agent Models of Organizational Intelligence

The Orchestration of Multi-Agent Systems: Architectures, Protocols, and Enterprise Adoption

MARO: Learning Stronger Reasoning from Social Interaction

LSTM-MAS: A Long Short-Term Memory Inspired Multi-Agent System for Long-Context Understanding

Do We Always Need Query-Level Workflows? Rethinking Agentic Workflow Generation for Multi-Agent Systems

Learning Latency-Aware Orchestration for Parallel Multi-Agent Systems

TopoDIM: One-shot Topology Generation of Diverse Interaction Modes for Multi-Agent Systems

Beyond Rule-Based Workflows: An Information-Flow-Orchestrated Multi-Agents Paradigm via A2A Communication from CORAL

LLM-Based Agentic Systems for Software Engineering: Challenges and Opportunities

Collaborative Multi-Agent Test-Time Reinforcement Learning for Reasoning

The End of Reward Engineering: How LLMs Are Redefining Multi-Agent Coordination

A Large-Scale Study on the Development and Issues of Multi-Agent AI Systems

StackPlanner: A Centralized Hierarchical Multi-Agent System with Task-Experience Memory Management

CTHA: Constrained Temporal Hierarchical Architecture for Stable Multi-Agent LLM Systems

DynaDebate: Breaking Homogeneity in Multi-Agent Debate with Dynamic Path Generation

Demystifying Multi-Agent Debate: The Role of Confidence and Diversity

Orchestrating Intelligence: Confidence-Aware Routing for Multi-Agent Collaboration

Belief in Authority: Impact of Authority in Multi-Agent Evaluation Framework

When Single-Agent with Skills Replace Multi-Agent Systems and When They Fail

ResMAS: Resilience Optimization in LLM-based Multi-Agent Systems

TCAndon-Router: Adaptive Reasoning Router for Multi-Agent Collaboration

When Numbers Start Talking: Implicit Numerical Coordination Among LLM-Based Agents

Bayesian Orchestration of Multi-LLM Agents for Cost-Aware Sequential Decision-Making

OptimAI: Optimization from Natural Language Using LLM-Powered AI Agents

Memory & RAG (57)

Corpus2Skill: Don't Retrieve, Navigate — Distilling Enterprise Knowledge into Navigable Agent Skills for QA and RAG

Semantic Level of Detail for Knowledge Graphs: Discovering Abstraction Boundaries via Spectral Heat Diffusion

BudgetMem: Learning Query-Aware Budget-Tier Routing for Runtime Agent Memory

Learning to Share: Selective Memory for Efficient Parallel Agentic Systems

CompactRAG: Reducing LLM Calls and Token Overhead in Multi-Hop Question Answering

Mitigating Hallucination in Financial Retrieval-Augmented Generation via Fine-Grained Knowledge Verification

Graph-based Agent Memory: Taxonomy, Techniques, and Applications

AI Agent Systems for Supply Chains: Structured Decision Prompts and Memory Retrieval

SOPRAG: Multi-view Graph Experts Retrieval for Industrial Standard Operating Procedures

ProcMEM: Learning Reusable Procedural Memory from Experience via Non-Parametric PPO for LLM Agents

Aggregation Queries over Unstructured Text: Benchmark and Agentic Method

DIVERGE: Diversity-Enhanced RAG for Open-Ended Information Seeking

JADE: Bridging the Strategic-Operational Gap in Dynamic Agentic RAG

ProRAG: Process-Supervised Reinforcement Learning for Retrieval-Augmented Generation

E-mem: Multi-agent based Episodic Context Reconstruction for LLM Agent Memory

ShardMemo: Masked MoE Routing for Sharded Agentic LLM Memory

When should I search more: Adaptive Complex Query Optimization with Reinforcement Learning

A2RAG: Adaptive Agentic Graph Retrieval for Cost-Aware and Reliable Reasoning

MemCtrl: Using MLLMs as Active Memory Controllers on Embodied Agents

AMA: Adaptive Memory via Multi-Agent Collaboration

When Iterative RAG Beats Ideal Evidence: A Diagnostic Study in Scientific Multi-hop Question Answering

Dep-Search: Learning Dependency-Aware Reasoning Traces with Persistent Memory

FadeMem: Biologically-Inspired Forgetting for Efficient Agent Memory

FastInsight: Fast and Insightful Retrieval via Fusion Operators for Graph RAG

Less is More for RAG: Information Gain Pruning for Generator-Aligned Reranking and Evidence Selection

DeepEra: A Deep Evidence Reranking Agent for Scientific Retrieval-Augmented Generated Question Answering

SPARC-RAG: Adaptive Sequential-Parallel Scaling with Context Management for Retrieval-Augmented Generation

Incorporating Q&A Nuggets into Retrieval-Augmented Generation

Augmenting Question Answering with A Hybrid RAG Approach

Utilizing Metadata for Better Retrieval-Augmented Generation

Deep GraphRAG: A Balanced Approach to Hierarchical Retrieval and Adaptive Integration

Grounding Agent Memory in Contextual Intent

Structure and Diversity Aware Context Bubble Construction for Enterprise Retrieval Augmented Systems

Topo-RAG: Topology-aware retrieval for hybrid text-table documents

Continuum Memory Architectures for Long-Horizon LLM Agents

Rethinking Memory Mechanisms of Foundation Agents in the Second Half: A Survey

The AI Hippocampus: How Far are We From Human Memory?

AtomMem: Learnable Dynamic Agentic Memory with Atomic Memory Operation

OpenDecoder: Open LLM Decoding to Incorporate Document Quality in RAG

Reliable Graph-RAG for Codebases: AST-Derived Graphs vs LLM-Extracted Knowledge Graphs

To Retrieve or To Think? An Agentic Approach for Context Evolution

Parallel Context-of-Experts Decoding for Retrieval Augmented Generation

SwiftMem: Fast Agentic Memory via Query-aware Indexing

Learning How to Remember: A Meta-Cognitive Management Method for Structured and Transferable Agent Memory

Beyond Dialogue Time: Temporal Semantic Memory for Personalized LLM Agents

Active Context Compression: Autonomous Memory Management in LLM Agents

Relink: Constructing Query-Driven Evidence Graph On-the-Fly for GraphRAG

Seeing through the Conflict: Transparent Knowledge Conflict Handling in RAG

CIRAG: Construction-Integration Retrieval and Adaptive Generation for Multi-hop Question Answering

Amory: Building Coherent Narrative-Driven Agent Memory through Agentic Reasoning

L-RAG: Balancing Context and Retrieval with Entropy-Based Lazy Loading

PRISMA: Reinforcement Learning Guided Two-Stage Policy Optimization in Multi-Agent Architecture for Open-Domain…

Controllable Memory Usage: Balancing Anchoring and Innovation in Long-Term Human-Agent Interaction

Beyond Static Summarization: Proactive Memory Extraction for LLM Agents

Membox: Weaving Topic Continuity into Long-Range Memory for LLM Agents

MAGMA: A Multi-Graph based Agentic Memory Architecture

HiMeS: Hippocampus-inspired Memory System for Personalized AI Assistants

SimpleMem: Efficient Lifelong Memory for LLM Agents

Eval & Observability (81)

Why Do AI Agents Break Rules? How Framing, Context, and Social Signals Shape Compliance

PerspectiveGap: A Benchmark for Multi-Agent Orchestration Prompting

RewardHarness: Self-Evolving Agentic Post-Training

ClawBench: Evaluating Browser Agents on Live Production Websites with Submission-Interception

StructEval: Benchmarking LLMs' Capabilities to Generate Structural Outputs

From Features to Actions: Explainability in Traditional and Agentic AI Systems

Agentic Uncertainty Reveals Agentic Overconfidence

AIRS-Bench: a Suite of Tasks for Frontier AI Research Science Agents

JADE: Expert-Grounded Dynamic Evaluation for Open-Ended Professional Tasks

Completing Missing Annotation: Multi-Agent Debate for Accurate Relevant Assessment

TrajAD: Trajectory Anomaly Detection for Trustworthy LLM Agents

Emulating Aggregate Human Choice Behavior and Biases with GPT Conversational Agents

Capture the Flags: Family-Based Evaluation of Agentic LLMs

PieArena: Frontier Language Agents Achieve MBA-Level Negotiation

ES-MemEval: Benchmarking Conversational Agents on Personalized Long-Term Emotional Support

HumanStudy-Bench: Towards AI Agent Design for Participant Simulation

Benchmarking Agents in Insurance Underwriting Environments

TriCEGAR: A Trace-Driven Abstraction Mechanism for Agentic AI

Sifting the Noise: A Comparative Study of LLM Agents in Vulnerability False Positive Filtering

Why Are AI Agent Involved Pull Requests (Fix-Related) Remain Unmerged? An Empirical Study

JAF: Judge Agent Forest

Stalled, Biased, and Confused: Uncovering Reasoning Failures in LLMs for Cloud-Based Root Cause Analysis

CAR-bench: Evaluating the Consistency and Limit-Awareness of LLM Agents under Real-World Uncertainty

More Code, Less Reuse: Investigating Code Quality and Reviewer Sentiment towards AI-generated Pull Requests

The Quiet Contributions: Insights into AI-Generated Silent Pull Requests

Agent Benchmarks Fail Public Sector Requirements

Interpreting Emergent Extreme Events in Multi-Agent Systems

Who Writes the Docs in SE 3.0? Agent vs. Human Documentation Pull Requests

Are We All Using Agents the Same Way? An Empirical Study of Core and Peripheral Developers Use of Coding Agents

DevOps-Gym: Benchmarking AI Agents in Software DevOps Cycle

Toward Architecture-Aware Evaluation Metrics for LLM Agents

Balancing Sustainability And Performance: The Role Of Small-Scale LLMs In Agentic AI Systems

Understanding Dominant Themes in Reviewing Agentic AI-authored Code

Let's Make Every Pull Request Meaningful: An Empirical Analysis of Developer and Agentic Pull Requests

Automated Structural Testing of LLM-Based Agents: Methods, Framework, and Case Studies

When AI Agents Touch CI/CD Configurations: Frequency and Success

Fingerprinting AI Coding Agents on GitHub

Interpreting Agentic Systems: Beyond Model Explanations to System-Level Accountability

AI builds, We Analyze: An Empirical Study of AI-Generated Build Code Quality

Will It Survive? Deciphering the Fate of AI-Generated Code in Open Source

LUMINA: Long-horizon Understanding for Multi-turn Interactive Agents

When Agents Fail to Act: A Diagnostic Framework for Tool Invocation Reliability in Multi-Agent LLM Systems

Agentic Confidence Calibration

Improving Methodologies for Agentic Evaluations Across Domains: Leakage of Sensitive Information, Fraud and…

MiRAGE: A Multiagent Framework for Generating Multimodal Multihop Question-Answer Dataset for RAG Evaluation

When Agents Fail: A Comprehensive Study of Bugs in LLM Agents with Automated Labeling

The Why Behind the Action: Unveiling Internal Drivers via Agentic Attribution

Tokenomics: Quantifying Where Tokens Are Used in Agentic Software Engineering

APEX-Agents

DRIFT: Detecting Representational Inconsistencies for Factual Truthfulness

CooperBench: Why Coding Agents Cannot be Your Teammates Yet

Insider Knowledge: How Much Can RAG Systems Gain from Evaluation Secrets?

Replayable Financial Agents: A Determinism-Faithfulness Assurance Harness for Tool-Using LLM Agents

AEMA: Verifiable Evaluation Framework for Trustworthy and Controlled Agentic LLM Systems

Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

ATOD: An Evaluation Framework and Benchmark for Agentic Task-Oriented Dialogue Systems

What Do LLM Agents Know About Their World? Task2Quiz

The Hierarchy of Agentic Capabilities: Evaluating Frontier Models on Realistic RL Environments

ViDoRe V3: A Comprehensive Evaluation of RAG in Complex Real-World Scenarios

M3-BENCH: Process-Aware Evaluation of LLM Agents Social Behaviors in Mixed-Motive Games

Mem2ActBench: A Benchmark for Evaluating Long-Term Memory Utilization in Task-Oriented Autonomous Agents

Active Evaluation of General Agents: Problem Definition and Comparison of Baseline Algorithms

VirtualEnv: A Platform for Embodied AI Research

FROAV: A Framework for RAG Observation and Agent Verification

Lost in the Noise: How Reasoning Models Fail with Contextual Distractors

RealMem: Benchmarking LLMs in Real-World Memory-Driven Interaction

IDRBench: Interactive Deep Research Benchmark

ToolGym: an Open-world Tool-using Environment for Scalable Agent Testing and Data Curation

TowerMind: A Tower Defence Game Learning Environment and Benchmark for LLM as Agents

MineNPC-Task: Task Suite for Memory-Aware Minecraft Agents

Internal Representations as Indicators of Hallucinations in Agent Tool Selection

Agent-as-a-Judge

Arabic Prompts with English Tools: A Benchmark

Effects of Personality Steering on Cooperative Behavior in LLM Agents

Analyzing Message-Code Inconsistency in AI Coding Agent-Authored Pull Requests

GUITester: Enabling GUI Agents for Exploratory Defect Discovery

Agent Drift: Quantifying Behavioral Degradation in Multi-Agent LLM Systems

M3MAD-Bench: Are Multi-Agent Debates Really Effective Across Domains and Modalities?

Why LLMs Aren't Scientists Yet: Lessons from Four Autonomous Research Attempts

LongDA: Benchmarking LLM Agents for Long-Document Data Analysis

The Rise of Agentic Testing: Multi-Agent Systems for Robust Software Quality Assurance

Project Ariadne: A Structural Causal Framework for Auditing Faithfulness in LLM Agents

ReliabilityBench: Evaluating LLM Agent Reliability Under Production-Like Stress Conditions

MAESTRO: Multi-Agent Evaluation Suite for Testing, Reliability, and Observability

Beyond Perfect APIs: WildAGTEval

Agent Tooling (98)

Steer, Don't Solve: Training Small Critic Models for Large Code Agents

Training Agents to Evolve with Their Harness: TaoLive Digital Avatar Agent Technical Report

Graph of States: Solving Abductive Tasks with Large Language Models

SCATE: Learning to Supervise Coding Agents for Cost-Effective Test Generation

Ouroboros: A Self-Developing Frontier Coding Agent with Reviewed Core Evolution

On Effectiveness and Efficiency of Agentic Tool-calling and RL Training

TraceCoder: A Trace-Driven Multi-Agent Framework for Automated Debugging

Generative Ontology: When Structured Knowledge Learns to Create

Structured Context Engineering for File-Native Agentic Systems

ProAct: Agentic Lookahead in Interactive Environments

Autonomous Question Formation for Large Language Model-Driven AI Systems

From Perception to Action: Spatial AI Agents and World Models

World Models as an Intermediary between Agents and the Real World

Engineering AI Agents for Clinical Workflows: A Case Study in Architecture, MLOps, and Governance

Autonomous Data Processing using Meta-Agents

MEnvAgent: Scalable Polyglot Environment Construction for Verifiable Software Engineering

Learning with Challenges: Adaptive Difficulty-Aware Data Generation for Mobile GUI Agent Training

AutoRefine: From Trajectories to Reusable Expertise for Continual LLM Agent Refinement

ToolTok: Tool Tokenization for Efficient and Generalizable GUI Agents

From Self-Evolving Synthetic Data to Verifiable-Reward RL: Post-Training Multi-turn Interactive Tool-Using Agents

Why Reasoning Fails to Plan: A Planning-Centric Analysis of Long-Horizon Decision Making in LLM Agents

SWE-Replay: Efficient Test-Time Scaling for Software Engineering Agents

Optimizing Agentic Workflows using Meta-tools

astra-langchain4j: Experiences Combining LLMs and Agent Programming

Meta Context Engineering via Agentic Skill Evolution

DataCross: A Unified Benchmark and Agent Framework for Cross-Modal Heterogeneous Data Analysis

CovAgent: Overcoming the 30% Curse of Mobile Application Coverage with Agentic AI and Dynamic Instrumentation

CUA-Skill: Develop Skills for Computer Using Agent

Textual Equilibrium Propagation for Deep Compound AI Systems

Should I Have Expressed a Different Intent? Counterfactual Generation for LLM-Based Autonomous Control

Insight Agents: An LLM-Based Multi-Agent System for Data Insights

Agentic Design Patterns: A System-Theoretic Framework

A Practical Guide to Agentic AI Transition in Organizations

JitRL: Just-In-Time Reinforcement Learning for Continual Learning in LLM Agents Without Gradient Updates

Think-Augmented Function Calling: Improving LLM Parameter Accuracy Through Embedded Reasoning

Paying Less Generalization Tax: A Cross-Domain Generalization Study of RL Training for LLM Agents

Think Locally, Explain Globally: Graph-Guided LLM Investigations via Local Reasoning and Belief Propagation

AI Agent for Reverse-Engineering Legacy Finite-Difference Code

PatchIsland: Orchestration of LLM Agents for Continuous Vulnerability Repair

DALIA: Towards a Declarative Agentic Layer for Intelligent Agents in MCP-Based Server Ecosystems

SWE-Pruner: Self-Adaptive Context Pruning for Coding Agents

REprompt: Prompt Generation for Intelligent Software Development Guided by Requirements Engineering

EvoConfig: Self-Evolving Multi-Agent Systems for Efficient Autonomous Environment Configuration

SemanticALLI: Caching Reasoning, Not Just Responses, in Agentic Systems

Controlling Long-Horizon Behavior in Language Model Agents with Explicit State Dynamics

Agentic Uncertainty Quantification

Agentic AI Governance and Lifecycle Management in Healthcare

Autonomous Business System via Neuro-symbolic AI

How to Build AI Agents by Augmenting LLMs with Codified Human Expert Domain Knowledge? A Software Engineering Framework

Agent Identity URI Scheme: Topology-Independent Naming and Capability-Based Discovery for Multi-Agent Systems

Toward Efficient Agents: Memory, Tool learning, and Planning

Toward self-coding information systems

A Lightweight Modular Framework for Constructing Autonomous Agents Driven by Large Language Models: Design,…

MagicGUI-RMS: A Multi-Agent Reward Model System for Self-Evolving GUI Agents via Automated Feedback Reflux

Agentic AI Meets Edge Computing in Autonomous UAV Swarms

Agentic Artificial Intelligence (AI): Architectures, Taxonomies, and Evaluation of Large Language Model Agents

Agentic Reasoning for Large Language Models

POLARIS: Typed Planning and Governed Execution for Agentic AI in Back-Office Automation

From Everything-is-a-File to Files-Are-All-You-Need: How Unix Philosophy Informs the Design of Agentic AI Systems

Towards AGI A Pragmatic Approach Towards Self Evolving Agent

EvoFSM: Controllable Self-Evolution for Deep Research with Finite State Machines

Investigating Tool-Memory Conflicts in Tool-Augmented LLMs

MAXS: Meta-Adaptive Exploration with LLM Agents

ToolACE-MCP: Generalizing History-Aware Routing from MCP Tools to the Agent Web

Beyond Single-Shot: Multi-step Tool Retrieval via Query Planning

OS-Symphony: A Holistic Framework for Robust and Generalist Computer-Using Agent

SAGE: Tool-Augmented LLM Task Solving Strategies in Scalable Multi-Agent Environments

Beyond Static Tools: Test-Time Tool Evolution for Scientific Reasoning

MegaFlow: Large-Scale Distributed Orchestration System for the Agentic Era

JudgeFlow: Agentic Workflow Optimization via Block Judge

R-LAM: Reproducibility-Constrained Large Action Models for Scientific Workflow Automation

OpenTinker: Separating Concerns in Agentic Reinforcement Learning

ARM: Role-Conditioned Neuron Transplantation for Training-Free Generalist LLM Agent Merging

PRISM: Disentangling SFT and RL Data via Gradient Concentration

ET-Agent: Incentivizing Effective Tool-Integrated Reasoning Agent via Behavior Calibration

No More Stale Feedback: Co-Evolving Critics for Open-World Agent Learning

CEDAR: Context Engineering for Agentic Data Science

ArenaRL: Scaling RL for Open-Ended Agents via Tournament-based Relative Ranking

Architecting AgentOps Needs CHANGE

Can We Predict Before Executing Machine Learning Agents?

EnvScaler: Scaling Tool-Interactive Environments for LLM Agent via Programmatic Synthesis

LIDL: LLM Integration Defect Localization via Knowledge Graph-Enhanced Multi-Agent Analysis

AT²PO: Agentic Turn-based Policy Optimization via Tree Search

M-ASK: Multi-Agent Search and Knowledge Optimization Framework

AgentDevel: Reframing Self-Evolving LLM Agents as Release Engineering

4D-ARE: 4-Dimensional Attribution-Driven Agent Requirements Engineering

XGrammar 2: Dynamic and Efficient Structured Generation Engine for Agentic LLMs

Transitive Expert Error and Routing Problems in Complex AI Systems

O-Researcher: An Open Ended Deep Research Model via Multi-Agent Distillation and Agentic RL

Architecting Agentic Communities using Design Patterns

SCRIBE: Structured Mid-Level Supervision for Tool-Using Language Models

Enhancing Model Context Protocol (MCP) with Context-Aware Server Collaboration

Enhancing LLM Instruction Following: An Evaluation-Driven Multi-Agentic Workflow for Prompt Instructions Optimization

InfiAgent: An Infinite-Horizon Framework for General-Purpose Autonomous Agents

The Path Ahead for Agentic AI: Challenges and Opportunities

AMER-RCL: Agentic Memory Enhanced Recursive Reasoning for Root Cause Localization in Microservices

Orchestral AI: A Framework for Agent Orchestration

AI Agent Systems: Architectures, Applications, and Evaluation

CaveAgent: Transforming LLMs into Stateful Runtime Operators

Actively Obtaining Environmental Feedback for Autonomous Action Evaluation Without Predefined Measurements

Warp-Cortex: An Asynchronous, Memory-Efficient Architecture for Million-Agent Cognitive Scaling on Consumer Hardware

DRACO: Fine-Grained Credit Assignment with Dynamic Rubrics for Long-Horizon Agent Training

AI Agent Security (83)

One Polluted Page Is Enough: Evaluating Web Content Pollution in LLM Recommenders

Internal Safety Collapse in Frontier Large Language Models

Confundo: Learning to Generate Robust Poison for Practical RAG Systems

Malicious Agent Skills in the Wild: A Large-Scale Security Empirical Study

Subgraph Reconstruction Attacks on Graph RAG Deployments with Practical Defenses

Zero-Trust Runtime Verification for Agentic Payment Protocols

Identifying Adversary Tactics and Techniques in Malware Binaries with an LLM Agent

Agent2Agent Threats in Safety-Critical LLM Assistants: A Human-Centric Taxonomy

Learning to Inject: Automated Prompt Injection via Reinforcement Learning

A Dual-Loop Agent Framework for Automated Vulnerability Reproduction

Human Society-Inspired Approaches to Agentic AI Security: The 4C Framework

MAGIC: A Co-Evolving Attacker-Defender Adversarial Game for Robust LLM Safety

TxRay: Agentic Postmortem of Live Blockchain Attacks

To Defend Against Cyber Attacks, We Must Teach AI Agents to Hack

SMCP: Secure Model Context Protocol

Persuasion Propagation in LLM Agents

When Agents "Misremember" Collectively: Exploring the Mandela Effect in LLM-based Multi-Agent Systems

"Someone Hid It": Query-Agnostic Black-Box Attacks on LLM-Based Retrieval

From Similarity to Vulnerability: Key Collision Attack on LLM Semantic Caching

TessPay: Verify-then-Pay Infrastructure for Trusted Agentic Commerce

Whispers of Wealth: Red-Teaming Google's Agent Payments Protocol via Prompt Injection

StepShield: When, Not Whether to Intervene on Rogue Agents

Delegation Without Living Governance

DRAINCODE: Stealthy Energy Consumption Attacks on Retrieval-Augmented Code Generation via Context Poisoning

Securing AI Agents in Cyber-Physical Systems: A Survey of Environmental Interactions, Deepfake Threats, and Defenses

Multimodal Multi-Agent Ransomware Analysis Using AutoGen

SHIELD: An Auto-Healing Agentic Defense Framework for LLM Resource Exhaustion Attacks

AgenticSCR: An Autonomous Agentic Secure Code Review for Immature Vulnerabilities Detection

AgentDoG: A Diagnostic Guardrail Framework for AI Agent Safety and Security

When Personalization Legitimizes Risks: Uncovering Safety Vulnerabilities in Personalized Dialogue Agents

Multi-Agent Collaborative Intrusion Detection for LAE-IoT

Faramesh: A Protocol-Agnostic Execution Control Plane for Autonomous Agent Systems

A Systemic Evaluation of Multimodal RAG Privacy

Breaking the Protocol: Security Analysis of the Model Context Protocol Specification

Prompt Injection Attacks on Agentic Coding Assistants: A Systematic Analysis

Connect the Dots: Knowledge Graph-Guided Crawler Attack on Retrieval-Augmented Generation Systems

Securing LLM-as-a-Service for Small Businesses: An Industry Case Study of a Distributed Chatbot Deployment Platform

Interoperable Architecture for Digital Identity Delegation for AI Agents with Blockchain Integration

INFA-Guard: Mitigating Malicious Propagation via Infection-Aware Safeguarding in LLM-Based Multi-Agent Systems

Query-Efficient Agentic Graph Extraction Attacks on GraphRAG Systems

NeuroFilter: Privacy Guardrails for Conversational LLM Agents

VirtualCrime: Evaluating Criminal Potential of Large Language Models via Sandbox Simulation

PINA: Prompt Injection Attack against Navigation Agents

Prompt Injection Mitigation with Agentic AI, Nested Learning, and AI Sustainability via Semantic Caching

CODE: A Contradiction-Based Deliberation Extension Framework for Overthinking Attacks on Retrieval-Augmented Generation

AgenTRIM: Tool Risk Mitigation for Agentic AI

Efficient Privacy-Preserving Retrieval Augmented Generation with Distance-Preserving Encryption

Taming Various Privilege Escalation in LLM-Based Agent Systems: A Mandatory Access Control Framework

Institutional AI: Governing LLM Collusion in Multi-Agent Cournot Markets via Public Governance Graphs

SD-RAG: A Prompt-Injection-Resilient Framework for Selective Disclosure in Retrieval-Augmented Generation

Beyond Max Tokens: Stealthy Resource Amplification via Tool Calling Chains in LLM Agents

Hidden-in-Plain-Text: A Benchmark for Social-Web Indirect Prompt Injection in RAG

Breaking Up with Normatively Monolithic Agency with GRACE: A Reason-Based Neuro-Symbolic Architecture for Safe and…

AgentGuardian: Learning Access Control Policies to Govern AI Agent Behavior

Agent Skills in the Wild: An Empirical Study of Security Vulnerabilities at Scale

CaMeLs Can Use Computers Too: System-level Security for Computer Use Agents

Blue Teaming Function-Calling Agents

Too Helpful to Be Safe: User-Mediated Attacks on Planning and Web-Use Agents

Semantic Laundering in AI Agent Architectures: Why Tool Boundaries Do Not Confer Epistemic Warrant

Towards Verifiably Safe Tool Use for LLM Agents

MCP-ITP: An Automated Framework for Implicit Tool Poisoning in MCP

Overcoming the Retrieval Barrier: Indirect Prompt Injection in the Wild for LLM Systems

MemTrust: A Zero-Trust Architecture for Unified AI Memory System

SafePro: Evaluating the Safety of Professional-Level AI Agents

Agentic LLMs as Powerful Deanonymizers: Re-identification of Participants in the Anthropic Interviewer Dataset

Toward Safe and Responsible AI Agents: A Three-Pillar Model for Transparency, Accountability, and Trustworthiness

VIGIL: Defending LLM Agents Against Tool Stream Injection via Verify-Before-Commit

Memory Poisoning Attack and Defense on Memory Based LLM-Agents

STELP: Secure Transpilation and Execution of LLM-Generated Programs

Conformity and Social Impact on AI Agents

Defense Against Indirect Prompt Injection via Tool Result Parsing

Autonomous Agents on Blockchains: Standards, Execution Models, and Trust Boundaries

BackdoorAgent: A Unified Framework for Backdoor Attacks on LLM-based Agents

HoneyTrap: Deceiving LLM Attackers with Resilient Multi-Agent Defense

SoK: Privacy Risks and Mitigations in Retrieval-Augmented Generation Systems

AgentMark: Utility-Preserving Behavioral Watermarking for Agents

Structural Representations for Cross-Attack Generalization in AI Agent Threat Detection

Lying with Truths: Open-Channel Multi-Agent Collusion for Belief Manipulation via Generative Montage

MCP-SandboxScan: WASM-based Secure Execution and Runtime Analysis for MCP Tools

Harm in AI-Driven Societies: An Audit of Toxicity Adoption on Chirper.ai

Trajectory Guard: A Lightweight, Sequence-Aware Model for Real-Time Anomaly Detection in Agentic AI

Mapping Human Anti-collusion Mechanisms to Multi-agent AI

Making Theft Useless: Adulteration-Based Protection of Proprietary Knowledge Graphs in GraphRAG Systems

When Agents See Humans as the Outgroup: Belief-Dependent Bias in LLM-Powered Agents

See category
94

Table of Contents

hesreallyhim/awesome-claude-code

A hand-picked collection of the finest of resources for the most awesome of agents, Claude Code, the undisputed champion of coding companions, from the unstoppable team…

Fresh★ 55k202 entriesPushed today
94

Awesome Agent Skills

VoltAgent/awesome-agent-skills

A curated collection of 1000+ agent skills from official dev teams and the community, compatible with Claude Code, Codex, Gemini CLI, Cursor, and more.

Fresh★ 35k839 entriesPushed today
93

Awesome Machine Learning

josephmisiti/awesome-machine-learning

A curated list of awesome Machine Learning frameworks, libraries and software.

Fresh★ 74k1188 entriesPushed 7 days ago
92

Awesome Production Machine Learning

EthicalML/awesome-production-machine-learning

A curated list of awesome open source libraries to deploy, monitor, version and scale your machine learning

Fresh★ 21k519 entriesPushed 3 days ago
92

AWESOME DATA SCIENCE

academic/awesome-datascience

:memo: An awesome Data Science repository to learn and apply for real world problems.

Fresh★ 30k881 entriesPushed today
91

Static Analysis

analysis-tools-dev/static-analysis

⚙️ A curated list of static analysis (SAST) tools and linters for all programming languages, config files, build tools, and more. The focus is on tools which improve…

Fresh★ 15k528 entriesPushed 8 days ago