Skip to content

Entry

AgentBench

Appears in 7 awesome lists

Multi-environment agent benchmark (OS, DB, web, code) with a structured eval pipeline. Worth studying for its environment isolation design and task definition format when building custom eval environments for your harness.

Open github.comthudm/agentbench

Found in these lists

When LLM Agents Meet Reinforcement Learning

Section: Environment · Tsinghua University

FreshScore 81

AI Game DevTools (AI-GDT)

Section: Game (World Model & Agent) · A Comprehensive Benchmark to Evaluate LLMs as Agents.

ActiveScore 77

Awesome Ai Agents 2026

Section: Benchmarks · 8-environment LLM agent benchmark.

ActiveScore 74

Awesome Autoresearch

Section: Evaluation & benchmarks · Comprehensive benchmark for LLM-as-Agent evaluation across 8 distinct environments. ICLR 2024.

FreshScore 86

awesome-ChatGPT-repositories

Section: NLP · A Comprehensive Benchmark to Evaluate LLMs as Agents

FreshScore 87

Awesome Harness Engineering

Section: Verification & CI Integration · Multi-environment agent benchmark (OS, DB, web, code) with a structured eval pipeline. Worth studying for its environment isolation design and task definition format when building custom eval environments for your harness.

FreshScore 88

Awesome AI Agents: Tools, Resources, and Projects

Section: Repositories · AgentBench v0.2 is a benchmark designed to evaluate Large Language Models as agents across a diverse set of environments, enhancing framework usability and extending model evaluations github

SlowScore 68

LangChain

Langchain integrates various providers like Anthropic, AWS, and OpenAI, and offers tools for components such as LLMs, chat models, and data analysis, supporting functionalities from Alpha Vantage to YouTube github | docs

In 20 listsDetails

LlamaIndex

(MIT) provides modules for structured outputs at different levels of abstraction, including output parsers for text completion endpoints, Pydantic programs for mapping prompts to structured outputs using function calling or output parsing, and pre-defined Pydantic programs for specific output types.

In 14 listsDetails

Dify

February 2026 release making human oversight a native workflow primitive: suspend execution at critical decision points, expose review-and-edit UI mid-flow, and route subsequent execution based on human action (approve/reject/escalate). Demonstrates how HITL transitions from bolt-on approval gates…

In 14 listsDetails

AutoGen

Microsoft's multi-agent conversation framework with a complete AgentChat layer covering agent loop, tool integration, termination conditions, and human-in-the-loop. The most comprehensive open-source reference for large-scale multi-agent harness design.

In 14 listsDetails

Pipecat

Handles frame management, streaming media coordination, and pipeline orchestration between ASR/LLM/TTS services for sub-800ms Total Turn-Around Time voice interactions. The missing harness primitive for voice agents: manages backpressure, handles frame queueing, and exposes a simple async…

In 7 listsDetails

OpenAgents

The web-browsing agent module of the OpenAgents platform (HKU). Enables autonomous navigation of websites via natural language, as part of a larger multi-modal agent framework.

In 7 listsDetails

ChatDev

ChatDev is a virtual software company utilizing intelligent agents to revolutionize the digital world through programming, offering a highly customizable framework and integrating innovative approaches like Experiential Co-Learning, Docker support, Git management, and Human-Agent Interaction…

In 6 listsDetails

SWE-agent

SWE-agent takes a GitHub issue and tries to automatically fix it, using GPT-4, or your LM of choice. It solves 12.29% of bugs in the SWE-bench evaluation set and takes just 1.5 minutes to run.

In 6 listsDetails