On the Evaluation of Agent Harnesses
Separating what the model contributes to an agent from what the system around it contributes.
In this essay Contents
Abstract
I spent a lot of the last year training LLM agents and building the harnesses that frontier models run inside, partly at AWS and partly in my own research. I kept notes along the way, and this essay turns them, checked against the published work, into a guide for anyone trying to find their way around this area and the tradeoffs that come with it. It covers what a harness actually is; when it pays to put effort into the harness and when a better model is the cheaper fix; how much structure to impose before it starts getting in the model’s way; when running many agents in parallel helps and when the coordination costs more than it buys; why sampling more only pays off if you can recognize a good answer; and why off-the-shelf agents are convenient to use but hard to experiment with. It also covers why agent benchmarks are so hard to read, and how I would set up an experiment that separates what the model contributes from what the harness contributes. The thread running through all of it is that an agent’s capability comes from the whole system, so many of the hard questions turn out to be questions about measurement. If you’re working on agents, I hope some of it saves you time.
Every few months the AI conversation swings between two camps. A new frontier model comes out, capability jumps, and people conclude that the model is the only thing that matters: a good enough model will work around mediocre scaffolding, and no harness will rescue weak weights. A few months later someone takes an existing model, gives it better tools, better context handling, a better execution environment, parallel workers and a verification loop, gets striking results, and the slogan becomes “the harness is the moat.” I have stopped finding either side convincing.
The trouble starts with treating an agent as a model. An agent is a system that runs at inference time. A model looks at a partly assembled state, picks an action through some interface, sees what happens, updates its working state, and keeps going until it decides to stop. Changing the interface alone can change measured capability by a lot. SWE-agent [1] replaced a generic Linux shell with an interface built for the agent and gained 10.7 percentage points on a 300-task SWE-bench ablation, with the same weights. Going the other way, SWE-Bench Pro [2] runs every model through a more standardized scaffold and still sees big differences between models; in the original study nothing evaluated got above 25% pass@1. BrowserGym [3] and AgentBench [4] tell a similar story. You can’t analyze the model or the harness in isolation, because what you measure comes from the pair.
I find it helpful to write this as (1) I wrote it as a product because the factors don’t add up independently. A harness that helps one model can get in another’s way. A tool abstraction that suits one kind of task can ruin another. Spending more compute at test time buys nothing if the system can’t tell its good trajectories from its bad ones. And a very strong model will look ordinary if the evaluation sandbox keeps killing its processes. When Anthropic varied only the infrastructure configuration on Terminal-Bench, scores moved by about 6 percentage points, and infrastructure failures dropped from 5.8% to 0.5% once the limits were loosened. [5] What I would like to know instead is
Which computation should live in the model, which should live in the harness, and when capability goes up, how do we tell which of the two was responsible?
I don’t think anyone has a good answer yet.
What exactly is a harness?
You can write a working agent in an afternoon. Hand a language model a goal and a few tools, ask it what to do next, run whatever it picks, paste the result into its context, and loop until it says it is done.
state = initialize(task)
while not done:
context = construct_context(task, state, history, tools)
action = model(context)
observation = environment.execute(action)
history.append(action, observation)
state = update(state, action, observation)
ReAct [6] gave this loop its best-known form, alternating reasoning steps with actions in the environment. With only one or two in-context examples, the original paper beat earlier imitation and RL baselines by 34 and 10 percentage points (absolute) on ALFWorld and WebShop. Reflexion [7] added verbal feedback and an episodic memory. Language Agent Tree Search [8] put Monte Carlo Tree Search and reflection on top of acting. Most agents I see today are some mix of these few ideas.1
That first version is misleadingly easy, because the difficulty is hidden inside construct_context, environment.execute, update and done. When I say agent harness, I mean everything that turns a foundation model into a stateful policy acting on some outside environment. In practice that covers the system and task prompts, how context gets assembled and cut down, which actions and tools exist, how tool calls are serialized, how the environment’s state is turned into observations, memory and retrieval, planning and decomposition, how results are checked and when to retry, who delegates to whom, concurrency and scheduling, permissions and sandboxing, when to stop, and how to pick or combine candidate answers. Anthropic has started calling this context engineering [11], by which they mean curating everything the model gets to see (tools, retrieved material, message history, memory, intermediate results), a much bigger job than polishing the wording of a prompt.2
In symbols, if a model with parameters picks action as (2) then it is choosing from , and never sees the world state itself. is assembled by the harness, roughly as (3) with the task, the memory, the tools as the model sees them, and whatever budget is left. So the harness decides both what the policy observes and what it is able to do. Seen from RL, swapping the harness can amount to swapping the environment the policy lives in. Agent Lightning [13] arrives at the same place from the training side. It models any agent run as an MDP and keeps rollout execution separate from credit assignment. This is why two systems built on the same model can differ so much in what they can do.
The scaffold can move the score by a lot
The cleanest evidence comes from software engineering, where you can usually find out whether the agent succeeded by running the code. SWE-bench [14] launched with 2,294 real GitHub issues across 12 Python repositories, and the best model tested at the time, Claude 2, resolved 1.96% of them. The SWE-agent authors [1] wondered what would happen if the computer interface were designed around a language model’s needs instead of a person’s. Their agent-computer interface (ACI) gave the model structured commands to search, view and edit a repository, and it phrased its feedback with the model in mind. Using GPT-4 Turbo, SWE-agent reached 12.47% on the full benchmark. In their controlled 300-instance ablation, the custom interface solved 10.7 percentage points more tasks than the same setup with a plain shell.
That gain came from changes made outside the model. The ACI changed which actions were cheap to take, how feedback read, how files were presented and how errors were explained. I think of it like changing an instruction set or an API. What can be computed in principle stays the same, but how hard it is to write down a good policy changes a great deal. The SWE-agent authors argued that language models are a new kind of computer user, and that interfaces designed for human cognition may not suit them. [1, 14]
Agentless [15] points in nearly the opposite direction. It throws out most of the autonomous loop and runs three fixed stages: find the faulty code, write a repair, validate it. The LLM never gets to choose what to do next. On SWE-bench Lite it fixed 96 of 300 issues (32.0%) at a reported average model cost of $0.70 per problem, ahead of every open-source autonomous agent in the paper’s comparison. That fits the argument in AI Agents That Matter [16], which asks for agent designs to be compared on cost and performance together, since success rate alone hides the price. I don’t read these as conflicting results. SWE-agent shows that a better interface helps an autonomous agent. Agentless shows that you shouldn’t assume autonomy is helping at all until you’ve checked. In both cases what moved the number was how inference was structured.
People have also started letting search do the design. AFlow [17] writes agent workflows as code and uses MCTS to search over the resulting graphs. Across six reasoning and coding benchmarks its authors report an average 5.7% improvement over state-of-the-art manually designed workflows, with some configurations letting smaller models beat GPT-4o for a fraction of the inference cost. Automated Design of Agentic Systems [18] searches over prompts, control flow and model calls in a similar way, and finds that designs discovered on one model and domain carry over to others. Calling all this “glue code” undersells it. In some settings, building a harness is an optimization problem over a huge, discrete space of programs, solved at inference time.3
But the model still matters
I am not claiming a clever loop will turn any model into a frontier agent. Fix the scaffold and the differences between models are large and easy to see, in browser benchmarks and coding benchmarks alike. AgentBench [4] put 27 proprietary and open models across eight interactive environments and found wide gaps in long-horizon reasoning, decision-making and following instructions. SWE-Bench Pro was designed to cut down scaffold variance, and its model-to-model gaps are still big. Standardizing makes the model’s contribution easier to see without making it any smaller. [4, 2]
Anthropic’s production research system offers a nice breakdown. On BrowseComp-style research tasks, the number of tokens used accounted for roughly 80% of the variance in performance, and tokens, tool calls and choice of model together accounted for about 95%. Yet moving the subagents from Claude Sonnet 3.7 to Sonnet 4 helped more than giving Sonnet 3.7 twice the token budget. The inference setup and the model were both doing real work. [23]
So I think of the boundary between model and harness as an iso-performance frontier. To hit some target capability there may be many workable combinations (4) where stands for model capability, for how sophisticated the harness is, and for test-time compute. With a stronger model you may get away with less search. A weaker model can sometimes close the gap with more sampling and checking. Good tools save context, and a strong verifier is what makes heavy sampling pay off. How much of one you can trade for another depends on the task, and the exchange rates are far from linear. The mistake is to assume one of the three always dominates.4
The prototype-production discontinuity
Harness work is easy for the first week and miserable after that. I think the reason is that most of what makes it hard costs almost nothing at small scale. One agent working for ten minutes barely needs a scheduler. Its whole state fits in the context window, so there’s nothing to merge. A failed API call can just be retried. Every tool schema fits in the prompt. Logging is a print statement, and nobody minds if the agent spends half a minute looking in the wrong directory. Once tasks run for hours, every one of those conveniences disappears.
Anthropic tried running capable models on long coding projects [25] and found that letting them continue across context compactions wasn’t enough. The agents bit off too much at once, hit context limits with the code in a half-broken state, forgot what earlier sessions had done, or announced that the project was finished when it wasn’t. The harness that fixed this added progress files that persist between sessions, a structured feature list, git history, setup scripts, a small goal for each session and a written handoff at the end of every session.5
A longer context window doesn’t make the problem go away. Lost in the Middle [27] showed that models with long contexts use them unevenly, and do noticeably worse when the evidence they need is buried in the middle. So agent builders have moved toward managing context like a cache instead of letting it grow as a log. They summarize, keep state in files outside the context, retrieve only what is needed and give specialist subagents a clean window. In Anthropic’s research system a subagent may burn tens of thousands of tokens on its own work and send only 1,000–2,000 high-value tokens back to the coordinator. [27, 23]
Tools hit the same wall. Five tools can go straight into the system prompt, but five hundred can’t. Anthropic describes setups where tool definitions and tool output used up 50,000+ tokens before the agent has meaningfully processed the user’s task, which pushed them toward discovering tools on demand and calling tools from code. A tool’s description also works as a prompt. Its namespace, what its arguments mean, the shape of its output and the wording of its error messages all change how the model behaves. [28]6 That’s how a demo harness can be 100 lines long while a production one grows into something like an operating system.
Prompt engineering becomes context engineering
Prompting feels like the dull part of the stack until a long rollout fails with no error, and you eventually work out that the model was handed the problem at the wrong level of abstraction. Models are still very sensitive to wording. OPRO [29] let an optimizer write instructions and beat human-written ones by up to 8% on GSM8K and 50% on selected BIG-Bench Hard tasks. DSPy [30] compiled LM pipelines automatically and did much better than hand-built few-shot prompts in its case studies. Experiments that move prompts between model families find that instructions don’t transfer cleanly, so the best phrasing for one model may do nothing for another, or make it worse. [29, 30]7
So I’ve stopped looking for magic sentences. I think of the prompt as part of the observation function, one of the things that shapes the policy the model ends up running. A context policy chooses (5) within a fixed token budget, where is the instructions, the goal of the task, the current state, whatever was retrieved, memory, demonstrations and the tool definitions. The harness should aim for the most decision-relevant information per token, which is usually far less context than it could fit.
Tell a model “solve this issue” and it has to work out how to break the problem down, how the repository is laid out, whether the tests are good enough and when to stop, all on its own. Scripting every step can go just as badly, because it locks out approaches the model might have found itself. Anthropic’s word for getting this right is “altitude”: specific enough to steer, loose enough that you aren’t hardcoding reasoning that will break. [11]8 I suspect this explains a lot of the variation between models in agent systems that people don’t notice. A harness carries assumptions about how its model reasons. Change the model and some of those assumptions stop being true. Anthropic has said nearly the same thing: harness assumptions go stale as models get better [35].
Tools are part of the model’s effective action space
The comparison that helps me most here comes from compilers. In principle a CPU can compute almost anything, but its instruction set decides what is convenient to write. In the same way, a language model can in principle type any shell command, and the harness decides which useful changes to the world it can make easily and reliably. Compare giving the agent raw shell tools,
grep
sed
awk
cat
python
git
...
with giving it semantic ones:
search_repository(query)
read_file(path, line_start, line_end)
apply_patch(diff)
run_tests(target)
The second set boxes the model in more and makes badly formed actions less likely. Whether that helps depends on which model you have. SWE-agent’s +10.7-point ACI result is a case where it clearly helped. [1] I’d score a tool on at least four things: (6) Affordance is whether the actions line up with the plans the model would want to carry out. Discoverability is whether the model can find the right tool when there are dozens or thousands. Feedback is whether a failed call tells the model enough to try something else. Token cost is whether calling the tool floods the context and crowds out the agent’s working memory.
Big outputs make all of this worse. A raw database dump, a full browser DOM, a compiler trace or a listing of the whole repository may contain everything, and still be close to useless to the model. More and more harness code exists to turn raw output from the environment into something the model can work with. SWE-agent’s ACI, the push to standardize browser-agent environments and dynamic tool discovery all come from this idea. In effect the harness designer is hand-crafting the model’s input representation, with no gradients to help. [1, 28]
Subagents turn harness engineering into distributed-systems engineering
A single agent is a loop. Run three hundred of them and you are operating infrastructure, and the hard problem stops being the prompt and becomes coordination. Say the parent task breaks into a dependency DAG (7) where each is one agent run, and an edge means one agent needs another’s output. The scheduler now has to decide when to start each agent, what state to hand it, how much compute to give it, whether a speculative branch is worth the cost and how to merge the results when they come back.
Code makes this harder because of the repository. Two agents editing one checkout will race each other. Giving each its own worktree or clone removes the contention but leaves a merge for later. Let agents spawn their own children and the scheduler becomes recursive. Shared memory needs synchronized writes, and without it every piece of information has to be passed along explicitly. Run things asynchronously and a result can come back to a parent whose state has already changed. Anyone who has built distributed systems knows the failure list: stragglers, stale state, conflicting writes, duplicated work, dead tasks, starvation, cascading failures, runaway fan-out, backpressure, costly barriers and processes that disagree about whether they’re done.
Anthropic’s multi-agent research system hit many of these with only a few workers. The lead agent usually starts 3–5 subagents in parallel, each of which runs several searches at once. That cut latency by up to 90% on complex research questions. Because the orchestration is synchronous, though, the coordinator still waits on the slowest subagent, can’t redirect a worker that is already running, and the workers have no good way to talk to each other. Anthropic says asynchronous execution would be appealing but is hard to get right, because of state consistency, coordination and the way errors spread. [23]9
Some experiments already go much further. In 2026 Anthropic ran 16 parallel agents over nearly 2,000 Claude Code sessions, spent about $20,000 on API calls, and ended up with a C compiler in Rust of roughly 100,000 lines that can build Linux 6.9 for x86, ARM and RISC-V. I don’t know whether that is economical as a product, but as a systems experiment it is striking. The model’s intelligence was one part of a bigger engineering problem that also included splitting up the work, maintaining a shared test suite, disciplined merging and keeping thousands of runs moving forward without a human. [36] With 1,000 or 10,000 workers, I doubt “multi-agent prompting” is still the right way to describe it. It looks more like a distributed inference operating system.
Multi-agent scaling is not automatically positive
The natural assumption is that if one agent helps, of them will help more. That holds only under certain assumptions about how their outputs get combined. If each of independent agents solves a task with probability , the chance that at least one succeeds is (8) At , one agent covers 10%, ten cover about 65% and a hundred cover more than 99.99%. Real agents aren’t independent, though, and having a correct trajectory somewhere in the pile doesn’t mean you can pick it out.
More Agents Is All You Need [37] got gains on several benchmarks just by sampling many answers and voting. Anthropic’s research setup, with Opus 4 coordinating Sonnet 4 workers, scored 90.2% higher on an internal research evaluation than Opus 4 working alone. It also used about 15× the tokens of a normal chat interaction, where a single agent uses about 4×. [37, 23]
How much you gain depends heavily on the task. Towards a Science of Scaling Agent Systems [38], a 2025 preprint, ran 180 configurations covering five coordination architectures and three model families. The results show nearly every way that just adding agents can backfire. On a financial-reasoning task that splits up easily, centralized coordination gave an 80.9% improvement. On sequential reasoning, the multi-agent versions did 39–70% worse. Once a single agent already scored above roughly 45%, adding coordination gave little or nothing, and sometimes cost performance. The authors also estimate that errors were amplified 17.2× without checks in one topology with independent agents, against 4.4× with a central coordinator.
Newer work gets to a similar place through information theory. Copies of the same worker quickly start making the same guesses, so beyond a point an extra copy mostly repeats evidence you already have. One 2026 study found that two diverse agents can match or outperform 16 homogeneous agents in the setups it tried. [38, 39] So the number that matters is probably not (9) but something more like (10) If ten copies share a model, a prompt, a retriever, a toolset and a blind spot, you have one researcher running ten times.10
Test-time compute is becoming a first-class scaling axis
This is where harness engineering meets the work on scaling inference. Self-consistency [41] was one of the first simple demonstrations: sample several chains of reasoning and take the most common answer. Over plain chain-of-thought it gained 17.9 percentage points on GSM8K, 11.0 on SVAMP, 12.2 on AQuA, 6.4 on StrategyQA and 3.9 on ARC-Challenge.
Repeated sampling pushes this further. Large Language Monkeys [42] varied the number of samples over four orders of magnitude. On SWE-bench Lite, the share of problems DeepSeek-Coder-V2-Instruct solved at least once went from 15.9% with a single sample to 56% with 250. The same paper also shows the catch. Without a verifier that can check answers objectively, majority voting and learned reward models level off, because the system can produce a correct solution without recognizing it. I want to state that difference explicitly: (11) Suppose I spend 1,000× the compute generating trajectories, and one of them is right while the other 999 are wrong in convincing ways. If I can’t reliably find the right one, most of what I paid for is wasted.
Tree of Thoughts [43], RAP [44] and LATS [8] try to spend inference compute more carefully. They score partial solutions along the way, where plain sampling would run every trajectory to the end on its own. RAP runs MCTS with the LM acting as both reasoner and world model. Using LLaMA-33B, it reported a 33% relative improvement over GPT-4 with chain-of-thought in one planning setting. LATS mixes MCTS, feedback from the environment and self-reflection, and reached 92.7% HumanEval pass@1 with GPT-4 in its original experiments. [44, 8]11
Broader studies of test-time scaling make the same point: where the compute goes matters. Scaling LLM Test-Time Compute Optimally [49] found that allocating compute adaptively was more than 4× as efficient as plain best-of- in its setup, and that with FLOPs held equal, a small model given extra test-time compute could sometimes beat a model 14× larger. That makes a harness, more and more, an inference-compute allocator. Each time it has another dollar to spend, it has to choose between another independent attempt, searching deeper on the current one, a critic, a verifier, a tool call, a new subagent, more context or a stronger model. I’d guess that choice ends up being the most important thing a harness does.
The optimal topology depends on the task
I don’t expect one agent architecture to win everywhere. Take a research question that is mostly breadth, (12) where the pieces barely depend on each other. Parallel workers are an obvious fit. Fixing a bug in a repository sits in between. Finding the fault, writing the fix and testing it all feed into each other, but you can sometimes pull them apart.
A proof, or any plan where each step builds on the last, (13) is a different story. Each step depends closely on the exact state the previous steps left behind. Parallel copies tend to duplicate each other or head off in incompatible directions, and stitching their output together can be harder than letting one agent carry on.
Anthropic reports that its multi-agent research setup does best on broad questions and is less obviously suited to coding work with tightly linked dependencies. Its 90.2% internal improvement shouldn’t be read as a multiplier that applies to any task. The controlled scaling studies agree: tasks that parallelize can gain a lot, and sequential ones can lose. [23, 38] One primitive I’d like future harnesses to have:
estimate_parallelizability(task)
→ (k, topology, budgets)
The harness wouldn’t fix “single agent” or “multi-agent” ahead of time. It would estimate at run time how tightly the task’s parts are coupled and pick a topology to match, much like a database query planner or a parallelizing compiler.
Harnesses are part of the post-training distribution
Things get more interesting once we train agents instead of only prompting them. When a model is trained inside an interactive loop, it doesn’t pick up some general skill called “tool use.” It learns a policy tied to one specific format for observations and one specific format for actions. If every training episode looks like
{"tool":"search","query":"..."}
and the deployed system suddenly wants
<browser.search q="..."/>
we’re relying on the model to generalize from the meaning, and there’s no guarantee it will do so fully. The longer the interactions, the more this matters. Search-R1 [50] uses reinforcement learning to teach models when and how to search, possibly several times, in the middle of reasoning. Across seven QA datasets it reports gains of 26% for Qwen2.5-7B, 21% for Qwen2.5-3B and 10% for Llama-3.2-3B over its baselines. What the model learns is bound up with the particular retrieval setup it practiced on.
AgentGym [51] goes after generalization by training agents across many environments instead of growing them inside one. I think its premise is right. Train in one interactive environment and you get a specialist, so broadly capable agents will need varied environments to learn in.
Agent Lightning [13] works on the plumbing. It tries to separate whatever the agent is doing from the RL trainer by putting a standard interface for transitions and credit assignment between them. We need something like this to tell whether training gave us a better model or only a better fit to one runtime.
What I’d want post-training to optimize is (14) averaged over a distribution of harnesses and tasks, instead of (15) with a single harness . Train against one harness and it becomes part of the environment, and the model may overfit to it without anyone noticing.12
The benchmark problem is worse than it looks
Agent benchmarks are much harder to read than static model benchmarks. A static QA benchmark is roughly (16) An agent benchmark scores the result of a random, closed-loop interaction, (17) and the number at the end depends on the model, the harness, how the environment was implemented, the compute and memory it was given, which tools were available, how the network behaved, when the run was cut off and how the grader works. I don’t trust small leaderboard gaps for this reason.
Anthropic tested this on Terminal-Bench 2.0. Same model, same harness, same tasks, and only the resource limits changed: scores varied by 6 points from the tightest configuration to the uncapped one, and infrastructure failures went from 5.8% to 0.5%. In a separate SWE-bench run on 227 problems with 10 trials each, giving the sandbox 5× the baseline RAM raised scores by 1.54 percentage points. [5]
Grading bugs can do even more damage. Anthropic describes an internal run of Claude Opus 4.5 on CORE-Bench that first scored 42%. After they fixed problems with the grader and loosened the scaffold, the same model scored 95%. Whatever you think of the model, a 53-point swing tells you an agent’s score is partly a measurement of the test setup. [58]13
SWE-bench has been through the same thing at the level of the whole benchmark. It started with 2,294 tasks. OpenAI and the SWE-bench authors then had 93 software developers review 1,699 tasks, found problems that were underspecified or had broken evaluation tests, and kept 500 as SWE-bench Verified. By February 2026 OpenAI was arguing that Verified had become contaminated too, to the point that top scores no longer measured software-engineering ability cleanly. They pointed to state-of-the-art results bunching between 74.9% and 80.9% over the previous six months, and to signs that every frontier model they tested had seen some of the benchmark’s material. [65, 66]
Moving to a newer benchmark doesn’t escape this. In July 2026 another OpenAI audit found serious problems in about 30% of SWE-Bench Pro tasks. Follow-up work on verified versions of these benchmarks argues that leaked rewards and badly specified tasks can change what we conclude about agents. [66] None of this is a knock on the people who build these benchmarks. Measuring agents is just very hard.
Reliability matters at least as much as mean success
Reporting only pass@1 is another mistake. Agents behave randomly. Picture one agent that solves each task 80% of the time, and another that also averages 80% across the benchmark but is erratic on any given task. You would not want to deploy them in the same places.
-bench [67] added pass to measure how often an agent succeeds on every one of repeated attempts. In its original retail experiments, the best function-calling agents succeeded less than 50% of the time on a single try, and pass fell below 25%. An agent that looks decent on one attempt can be much less dependable over repeated runs.
More recent work measures run-to-run randomness on the same task directly. A 2025 study used the intraclass correlation coefficient on agent evaluations and found ICC values from about 0.304 to 0.774 on GAIA, depending on the model. Its experiments suggest you need 8–16 trials for stable numbers on structured tasks and 32 or more on harder reasoning tasks. [68] This matters a lot when you compare harnesses. A change that takes accuracy from 55% to 58% but doubles the variance between runs can be worse in production, even though it wins on the leaderboard.
At a minimum, I would report (18) Collapsing all of that into one leaderboard number throws away too much.
How I would benchmark model versus harness
Say we have models and harnesses . The usual experiment compares two complete systems: (19) That can’t tell you whether the model or the harness made the difference. You need the crossed design, (20) filled in for as many pairs as you can afford, with the inference budget held equal.
For task , model , harness and repeat , model the outcome as (21) is the model main effect, the harness main effect, the model–harness interaction, how hard the task is, and the noise from sampling and execution.
My bet is that the interaction term will turn out large. Several results already point that way. Prompts that are optimized for one model behave differently on another. SWE-agent’s ACI carries over to other models, but not equally well. Automated workflows transfer with uneven gains. Multi-agent setups behave differently as the single agent gets stronger, and the effect of extra resources depends on how a model goes about the task. [1, 17, 38]
Then do it all again with budgets matched on each of (22) You have to match cost, because plenty of supposed architectural wins turn out to be extra inference spending. AI Agents That Matter [16] makes this complaint directly, and Anthropic’s research system is a concrete case, using about 15× normal chat tokens for its multi-agent mode. [16, 23] And publish the trajectories. With only a score, you have almost no way to find out why a run failed.
Ablate semantic components, not product names
Pitting Claude Code against Codex against OpenHands helps someone choosing a tool. It doesn’t teach us much. A benchmark meant to study harnesses should break them into parts: (23) for instance:
: prompt policy
: context-management policy
: action space
: observation representation
: memory
: scheduler/search strategy
: verifier
: delegation topology
: resource policy
Then change them one at a time: (24) That kind of isolation is what made SWE-agent’s interface study worth reading. Agentless was worth reading because it tested whether autonomy was earning its keep. AFlow and the other automated-design papers are interesting because they treat these choices as a space to search, instead of treating an agent framework as one thing you take or leave whole. [1, 15, 17] If harness engineering is going to become a science, I expect it to happen at this level of detail.14
Off-the-shelf agents are products, not neutral experimental instruments
Claude Code and Codex show how awkwardly agent harnesses sit between product and research tool. People use them because someone else has already done a lot of the tedious work: running tools, sandboxing, managing context, working with repositories, retrying, permissions and the user interface. If you run research on top of one, though, you take on all of its design decisions along with it.
When I wrote this, Claude Code’s public repository was released under an all-rights-reserved commercial license, so its internal runtime is not open source. OpenAI publishes the Codex CLI under Apache 2.0. For day-to-day use this hardly matters. For reproducible research it can, because whether you can read the orchestration logic, the prompts, the state transformations and the version history may decide whether two rollouts can be compared at all. [70, 71]
Post-training raises the stakes. If the rollout policy includes transformations you can’t see, you don’t fully know what policy generated the behavior. You can’t cleanly split credit between the model and orchestration choices you never saw, and you may not be able to rerun the same trajectories with a modified model. I have nothing against closed products. My point is that using the best available agent and running a controlled agent experiment are separate jobs, and one tool may not serve both.
OpenClaw is an instructive example of the other extreme
OpenClaw is a good counterexample to the belief that anything valuable has to come from the model weights. It doesn’t ship a new foundation model. It runs existing models behind a gateway that stays on all the time and connects to messaging apps, local machines and the user’s own tools. Persistence, memory, tool access, permissions, event triggers and being reachable from anywhere change what the system can do for its user, though none of it makes the underlying transformer any smarter. [72]
I wouldn’t take this as proof that the harness is the moat. Agents that run all the time with broad permissions also expose a much bigger attack surface, and security researchers have begun studying systems like it. [73] OpenClaw does show that the layer wrapped around inference can turn the same family of models into a very different product.
Long-horizon capability is partly a reliability problem
Part of the reason harnesses matter more as tasks get longer is arithmetic. If an agent must get important steps right and each succeeds independently with probability , a rough lower-bound estimate of overall success is (25) Even at , (26) (27) (28) Real steps aren’t independent, but the general lesson survives: a small error rate per step adds up to near-certain failure over a long enough run. That’s a big part of why long-running agents need ways to recover, check their work and save their progress, on top of whatever intelligence the model brings.
METR’s time-horizon work [74] measures something close to this: how long a task (timed by how long a human takes) an AI agent can complete with a given probability. Their original study put Claude 3.7 Sonnet’s 50% time horizon at about 50 minutes and found that the frontier horizon had doubled roughly every seven months since 2019. They credited the progress partly to better reliability, error recovery, reasoning and tool use, as well as to getting answers right the first time.
As the horizon grows, so does the payoff from recovery machinery. Halving a model’s mistakes can make it far more useful. A harness that catches and repairs half of the mistakes that remain can also improve end-to-end success by more than you’d expect. Because errors compound over steps, the two improvements multiply.
Verification may be more important than generation
If I had to pick one part of the harness to become much more important than the rest, I’d pick verification. A good generator can produce many plausible trajectories. The costly part is deciding which one to believe.
The inference-scaling papers show this over and over. Repeated sampling finds a correct answer somewhere in the batch very often in domains where answers can be checked, but you only get that benefit if the selector is good. Reflexion uses feedback from the environment to improve the next attempt. LATS steers its tree search with value estimates and real outcomes. Agentless generates several candidate patches and filters them with reproduction tests it writes itself plus the existing regression tests. [42, 7, 8, 15] So I’d split the probability of success into at least two pieces: (29) More test-time compute raises the first factor. Only better selection, which usually means a verifier, raises the second.
In software, theorem proving and math, the second factor can be cheap, because tests, proof checkers or exact answers tell you when you’re right. In research, product design and other open-ended work, picking the right candidate is a hard judgment call. That may be why brute-force scaling does so well on code and proofs and so much less cleanly in areas where nothing checks the answer.15
A good harness is an adaptive compute allocator
Putting this together, I’ve come to see the core job of a harness as online resource allocation. At each state it chooses among moves like (30) Each move costs and has some expected payoff, whether in information or in progress toward the goal. The ideal policy is roughly (31) Nobody can compute this exactly. It’s still a more useful target than “make the agent smarter.”
In practice: finish easy tasks fast, retrieve more when the task is ambiguous, sample more when answers can be checked, fan out when the task is broad, keep a single thread when it’s sequential, spend more on branches you’re unsure about, and kill branches that aren’t going anywhere. Adaptive test-time compute, Monte Carlo tree search, workflow optimization and multi-agent scheduling all fit inside this view. RAP, LATS, AFlow and compute-optimal inference each approximate the same problem in a different way: given a fixed amount of computation, how do you spread it over the space of possible reasoning paths? [44, 8, 17, 49]
What makes a harness good?
After building a fair number of these, I’d currently boil it down to six properties.
It exposes the right state
The model needs enough information to act well, but not so much that the useful parts get lost. Long-context studies, context-engineering practice and tool-interface work all favor choosing what to show over piling everything in. [27, 11]
It exposes the right actions
Good tools make useful actions easy to express, give helpful errors, and don’t make the model build everything out of awkward low-level steps. SWE-agent is the standard example. [1]
It externalizes durable state
Long tasks need checkpoints, saved artifacts, version history, structured memory and a way to pick up where they left off. Without those, the agent loses part of its memory at every context reset. [25]
It matches orchestration to task structure
Run things in parallel when they are conditionally independent. Throwing many agents at a task that is sequential at heart tends to raise both cost and error rate. [23, 38]
It has a verifier whenever possible
Generating more candidates only helps if you can recognize the good ones. [42]
It knows when to stop spending compute
Agents buy success with money and time. Anthropic’s advice is to use the simplest architecture that does the job, and Agentless and the cost-aware evaluation work show why that advice is sound. [78, 15, 16] None of these is surprising on its own. Getting all six right in one system is hard.
The hardest open problem is attribution
We are getting better at improving agents, but we are still bad at explaining why a given agent improved. Suppose system A beats system B by eight points. A might have a better model, a different prompt, more tokens, more tool calls, different tool definitions, better retrieval, a different observation format, more retries, a different verifier, more memory, more RAM, a longer timeout, more parallel workers, or a lucky overlap with contaminated benchmark data. Unless someone ran properly crossed ablations, “A is a better agent” tells you almost nothing.
Agent evaluations are especially exposed to this because the measuring equipment is part of the causal chain. The grader, sandbox, network, tool implementations and scheduler all change which trajectories can happen in the first place. The infrastructure-noise experiments, the contamination audits, the work on standard environments and the push for cost-aware evaluation all show the setup shaping the score. [5, 66, 16] So agent research should borrow from systems benchmarking more than from the way we have traditionally benchmarked models. Concretely, report the whole stack and version it, publish trajectories, fix the budgets, run multiple seeds, cross models with harnesses, measure variance and cost, and be skeptical of small gaps on leaderboards.16
The model–harness boundary will keep moving
Another reason this area is hard: every time models improve, some harness tricks stop being needed. When models couldn’t plan, harnesses supplied a planner. When they couldn’t remember, harnesses kept the memory. When their tool calls came out malformed, harnesses constrained the syntax, and when they couldn’t recover from errors, harnesses added retries. Each time a new generation learns one of these behaviors itself, the matching piece of outside machinery can go.
Better models also make room for new kinds of harness, though. Recursive subagents are pointless if the model can’t delegate reliably. Once it can, you suddenly have a whole distributed-systems design problem on your hands. So harness complexity doesn’t necessarily fall as models improve. It moves up a level of abstraction. Anthropic makes this argument in its writing on managed agents [35]. Every harness assumes certain things about what the model is bad at, and those assumptions need revisiting each time capabilities change. The engineer’s job shifts from writing each state transition by hand toward building an environment where the model can produce those transitions reliably on its own, which is roughly what I mean by harness engineering.17
From one agent to ten thousand
Right now the conversation is about giving one agent a shell. Before long it may be about scheduling 10,000 reasoning processes at once. If that happens, people who have worked on distributed compute will recognize the architecture: (32) Workers could run different models, with cheap ones exploring and expensive ones critiquing. Some might keep state for the whole job while others exist for a single tool call. The planner could widen the fan-out when it’s uncertain and narrow back to one trajectory when the work turns sequential. At that scale an “agent” looks less like a chatbot with tools and more like a compute graph generated at inference time.
AutoGen [83], MetaGPT [84] and similar frameworks were early experiments in this direction. Newer scaling studies have started asking which of these topologies actually help when budgets are controlled. [83, 84, 38] My guess is that this will be one of the main systems problems of the next few years.
So where does the intelligence live?
In many places at once. Pretraining and post-training compress some of it into the weights. More comes from the context assembled at inference time, the tools on offer and the way their outputs are presented, and from retrieval and external memory. A verifier, the scheduler, parallel sampling and the test suite each add some. And now and then, what looked like capability was really the benchmark configuration. That makes the whole compound system the thing worth studying: (33) Model research matters as much as it ever did. But once models are this capable, the architecture around them at inference time becomes a real source of capability too.
The literature I’ve gone through cuts both ways. Better interfaces have moved success rates by double digits without touching the weights. Automatically discovered workflows have beaten hand-designed ones. Repeated sampling has taken 15.9% one-shot coverage to 56%. A multi-agent research system beat a strong single agent by 90.2% on a broad internal evaluation. On the other side, Agentless beat more elaborate agents for pennies per problem, multi-agent setups made sequential reasoning 39–70% worse, infrastructure settings alone moved benchmark scores by six points, and coding evaluations that seemed solid were undermined by contamination and grading bugs. [1, 17, 42, 23, 38, 5, 66] Neither side of the old argument comes out of this list with a clean win. Taken together, these results say that agent intelligence is a property of the whole system. If that’s right, the hard work ahead is learning to build, tune and properly evaluate dynamic systems that contain foundation models but can’t be reduced to them.
The first version will still take an afternoon. Getting the last 20% working may be the whole field.
References
Notes
Cite this essay
Debargha Ganguly. “On the Evaluation of Agent Harnesses.” Research notes, October 1, 2026. https://blog.debargha.com/essays/on-agent-harnesses/
@misc{ganguly2026evaluationagent,
author = {Ganguly, Debargha},
title = {On the Evaluation of Agent Harnesses},
howpublished = {Research notes},
year = {2026},
month = oct,
url = {https://blog.debargha.com/essays/on-agent-harnesses/},
}