Skip to content

Entry

Harness Engineering

Appears in 4 awesome lists

OpenAI's framing of harness engineering as a discipline: how to design the scaffolding that lets Codex and similar agents operate reliably in an agent-first world.

Open openai.com

Found in these lists

Awesome Artificial Intelligence

Section: Software factories and agent orchestration · OpenAI's field report on building software with coding agents, repository constraints, automated checks, and human steering.

FreshScore 86

Awesome Harness Engineering

Section: Foundations · OpenAI's framing of harness engineering as a discipline: how to design the scaffolding that lets Codex and similar agents operate reliably in an agent-first world.

FreshScore 88

Awesome Prompts

Section: Harness Engineering · Official OpenAI post: "leveraging Codex in an agent-first world"

FreshScore 90

Best Of Agent Harnesses

Section: Related Resources · Environment design, intent, feedback loops, repo-as-system-of-record

FreshScore 84

ECC

Affaan Momin's agent-harness operating system: 68 specialized agents, 286 skills, hooks, memory, continuous learning, and AgentShield security scanning across Claude Code, Codex, Cursor, OpenCode, and other harnesses. The clearest open-source example of packaging an end-to-end engineering workflow…

In 7 listsDetails

langchain-ai/deepagents

LangChain's batteries-included agent harness (released April 2026) with built-in planning, filesystem tools, shell access, sub-agents, and auto-summarization. The clearest open-source demonstration of how a general-purpose coding agent harness can be made ready-to-run out of the box while…

In 7 listsDetails

Effective Harnesses for Long-Running Agents

Anthropic's pattern for maintaining agent progress across multiple context windows: an initializer agent sets up the environment once and hands off to a coding agent that makes incremental progress each session. The structured handoff mechanism — feature lists, git commits, and test gates as…

In 4 listsDetails

Mirage

Swaps the filesystem and bash providers for a mirage virtual workspace: file tools and shell commands run over mounted resources (RAM, S3, Redis, Slack, Gmail, Notion, Postgres) instead of the host disk, with per-mount read/write/exec modes, per-command sandbox routing (monty, pyodide, quickjs in…

In 3 lists

Loop Engineering

"Stop prompting. Design the loop." — practical patterns, starters & CLI (loop-audit, loop-init, loop-cost) for systems that discover work, hand it to agents, verify results, and persist state across Claude Code, Codex, Grok, and OpenCode; report-only week one, scores loops on a "Loop Ready" rubric…

In 3 lists

Demystifying Evals for AI Agents

Anthropic's framework for evaluating agent behavior: what to measure, how to build eval harnesses, and why unit-test-style evals fail for agents.

In 3 lists

Building Effective Agents

by Anthropic - Anthropic's foundational taxonomy of agent patterns — prompt chaining, routing, orchestrator-workers, and evaluator-optimizer — and when to use each.

In 2 lists

How We Built Our Multi-Agent Research System

by Anthropic - A practical account of orchestrator and subagent coordination, prompt design, and evaluation that maps directly to Claude Code's subagents and agent teams.

In 2 lists