Rethinking the Evaluation of Harness Evolution for Agents

Yike WangHuaisheng ZhuZhengyu HuYige YuanZhengyu ChenShakti SenthilHannaneh HajishirziYulia TsvetkovPradeep DasigiTeng Xiao

article2026arXiv39 citations

Demonstrates that automated agent scaffolding evolution fails to consistently outperform simple test-time scaling baselines or generalize to held-out tasks when evaluated under equal compute and feedback budgets.

Listen

Artificial intelligence agents driven by large language models increasingly rely on external harnesses—the supporting scaffolding of prompts, tools, memory, and control logic that structures how models interact with complex environments. While researchers have proposed automatic harness evolution to iteratively refine this scaffolding using benchmark feedback, existing evaluation protocols often test the evolved harnesses on the very same benchmarks used during optimization. This creates serious risks of task overfitting and obscures whether performance gains stem from genuine improvements in harness architecture or simply from spending extra search and inference compute during testing.

The article evaluates the true efficacy and generalizability of automatic harness evolution for AI agents. It investigates whether automatic harness modifications outperform standard, lightweight test-time scaling baselines under matched computational and feedback budgets.

The authors conducted a controlled empirical evaluation across 89 terminal coding tasks from the Terminal-Bench 2.1 benchmark using leading frontier models, including Claude Opus 4.6, GPT-5.4, and GPT-5.4 mini. They evaluated four distinct approaches under a unified budget framework: parallel sampling, sequential refinement, dataset-wide harness evolution, and instance-specific harness scaling. The experiments tested performance under three conditions: without access to unit test feedback, with access to unit test feedback, and across a strict data split where harnesses were trained, validated, and tested on completely disjoint task sets.

The findings show that automatic harness evolution fails to consistently outperform basic test-time scaling methods and exhibits severe overfitting. When unit test feedback is unavailable, automatic harness evolution averaged a 67.4% success rate, underperforming both direct baseline sampling at 68.2% and simple parallel sampling at 72.3%. When unit test feedback was provided, parallel sampling reached an 86.0% pass rate and sequential refinement achieved a 91.8% success rate across five attempts, substantially surpassing harness evolution's 75.8% and 86.2% respectively. Most critically, when evaluated on held-out tasks, harness evolution yielded virtually no benefit, providing an average gain of just 0.6 percentage points over the unevolved baseline.

These results indicate that automated harness revision primarily memorizes task-specific facts and fixes rather than discovering generalizable design strategies. The performance improvements reported in prior studies largely reflect the benefits of repeated sampling rather than superior scaffolding design. Consequently, allocating computational budgets directly to solution sampling and refinement is a far more reliable, cost-effective, and lower-risk strategy than dedicating resources to automated harness rewriting, which often introduces context bloat without solving fundamental reasoning failures.

Organizations developing AI agents should prioritize straightforward test-time scaling techniques, such as parallel candidate generation and sequential refinement, over complex automatic harness evolution. Research teams evaluating harness engineering must adopt rigorous protocols that strictly separate training tasks from evaluation tasks and benchmark all gains against compute-matched sampling baselines. Future work should investigate harness design on benchmarks where specialized tooling is strictly required to expand model capabilities.

The conclusions are supported by rigorous multi-model evaluations, but findings are bounded by the specific characteristics of the Terminal-Bench environment, where baseline models already achieve high performance and tasks require only minimal shell tool interfaces. Confidence in the core finding remains high: current automatic harness evolution methods risk severe overfitting and do not reliably deliver generalizable improvements over simpler inference-time search baselines.

  • Paper: Automated Design of Agentic Systems, Shengran Hu et al. (2025). Its automated search over agent designs provides a concrete precursor for understanding the harness-evolution methods whose generalizability this paper evaluates.
Cover for Rethinking the Evaluation of Harness Evolution for Agents

Abstract

We revisit the evaluation of automatic harness evolution for LLM agents. Existing harness evolution methods use unit test cases to search for harness configurations and then report final performance on the same public benchmark. This protocol raises two fundamental concerns. First, harness evolution is itself an iterative search procedure that repeatedly evaluates and revises candidate harnesses using task feedback. As in agentic test-time scaling, it should therefore be compared with simple task-level search baselines under matched feedback and inference budgets to determine whether its gains arise from improved harness design or from additional search alone. Second, because the search and the final evaluation share the same benchmark, the reported gains risk overfitting to that specific task set. To address these concerns, we conduct an extensive evaluation comparing harness evolution with simple test-time scaling and discovery baselines under comparable feedback and inference budgets, and also evaluate evolved harnesses on held-out tasks to assess whether the discovered improvements generalize. Experiments on Terminal-Bench 2.1 with GPT-5.4 and Claude Opus 4.6 show that automatic harness evolution does not consistently outperform simple test-time scaling methods and exhibits limited generalization. Our results raise important questions about the effectiveness of automatic harness evolution and highlight the need for fairer evaluation protocols and benchmarks for automatic harness design. Our code is available at this https URL.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Rethinking the Evaluation of Harness Evolution
  • 3.1 Preliminary
  • 3.2 Parallel Sampling
  • 3.3 Sequential Refinement
  • 3.4 Harness Evolution
  • 3.5 Harness Scaling
  • 4 Experiments
  • 4.1 Experimental Setup
  • 4.2 Without Unit Test Cases, Automatic Harness Evolution Underperforms Test-Time Scaling
  • 4.3 With Unit Test Cases, Automatic Harness Evolution Underperforms Test-Time Scaling
  • 4.4 Harness Evolution Fails to Generalize Beyond Training Tasks
  • 5 Discussion
  • 5.1 Rational Harness Edits but Marginal Gains
  • 5.2 Task Difficulty and Harness Sensitivity as Factors
  • 6 Conclusions
  • References
  • A Experimental Details
  • A.1 Initial Harness
  • A.2 Summarization Map
  • A.3 Agentic Harness Engineering (AHE)
  • A.4 Configuration
  • A.5 Evaluation Metrics
  • B Case Study

Knowls

  1. Knowl 1 — Matched-budget evaluation separates harness design from test-time search

    model/method

    The paper evaluates automatic harness evolution by comparing it with test-time methods under matched inference and feedback budgets. The comparison tracks three distinctions: whether a method receives unit-test outcomes, whether extra computation is spent on trajectories or on harness revisions, and whether updates are specific to one task or shared across tasks. Parallel sampling and sequential refinement spend computation on task trajectories; harness evolution revises a shared harness using evidence from a batch of tasks; harness scaling revises a harness for the particular task being solved. Evaluating both on the search tasks and on disjoint held-out tasks is intended to distinguish gains from reusable harness design from gains due to additional search or task-specific adaptation.

  2. Knowl 2 — Without unit-test feedback, parallel sampling beats harness evolution on average

    empirical result

    On Terminal-Bench 2.1 without unit-test feedback, the reported metric is pass@1, averaged over two independent runs. Scores for Claude Opus 4.6, GPT-5.4, GPT-5.4 mini, and their average, respectively, were: direct sampling with the initial harness, 69.9, 75.3, 59.4, 68.2; parallel sampling, 74.7, 79.2, 62.9, 72.3; sequential refinement, 73.0, 73.0, 61.8, 69.3; harness evolution, 71.4, 69.7, 61.3, 67.4; and harness scaling, 76.0, 78.1, 61.2, 71.8. Parallel sampling improved on the initial-harness average for all three models and had the highest average. Harness evolution scored below the initial-harness average, with its largest drop on GPT-5.4 (69.7 versus 75.3). Harness scaling achieved the highest individual score for Claude Opus 4.6, but did not exceed parallel sampling’s overall average.

  3. Knowl 3 — Unit-test feedback does not make harness methods outperform trajectory search

    empirical result

    On Terminal-Bench 2.1 with unit-test feedback, pass@1 measures average rollout success and pass@5 measures the fraction of tasks solved by at least one of five rollouts. Results are given as pass@1/pass@5 for Claude Opus 4.6, GPT-5.4, and the two-model average. Direct sampling with the initial harness scored 69.9/—, 75.9/—, and 72.9/—. Parallel sampling scored 84.8/84.8, 87.1/87.1, and 86.0/86.0; sequential refinement scored 83.1/90.4, 85.4/93.3, and 84.3/91.8; harness evolution scored 73.0/83.2, 78.6/89.3, and 75.8/86.2; harness scaling scored 83.1/89.9, 82.0/88.8, and 82.6/89.3. Thus, parallel sampling had the highest average pass@1, and sequential refinement had the highest average pass@5. Neither harness method exceeded the test-time scaling methods on either metric.

  4. Knowl 4 — Evolved harnesses yield little improvement on held-out tasks

    empirical result

    To test transfer beyond the search tasks, Terminal-Bench 2.1 was divided into 45 training tasks, 10 validation tasks, and 34 held-out test tasks. Harness evolution used unit-test feedback on the training tasks; the harness with the best validation performance was then evaluated on the held-out test tasks using pass@1. On the test tasks, the initial harness scored 63.3 with Claude Opus 4.6 and 72.1 with GPT-5.4, averaging 67.7. The selected evolved harness scored 64.5 (+1.2), 72.1 (+0.0), and 68.3 (+0.6), respectively. The reported gains were therefore small and absent for GPT-5.4, indicating limited transfer from the search tasks in this evaluation.

  5. Knowl 5 — Harness evolution searches for a shared harness using task-batch evidence

    algorithm

    Harness evolution aims to produce a reusable harness by revising it across a batch of tasks. Starting from an initial harness, each iteration runs the agent on the task batch using the current harness and records the harness and resulting trajectories in an experience store. When unit tests are available, the store also records each rollout’s pass/fail outcome. A meta agent receives a summarized view of the accumulated evidence and proposes the next harness. Without unit tests, the final harness is the latest one; with unit tests, the selected harness is the one with the highest aggregate success rate across the task batch or a validation set. The agent then runs on the target task using that selected harness. In the experiments, harness evolution was instantiated with Agentic Harness Engineering’s propose-and-refine loop, with its external harness-retrieval component disabled so that the method relied on collected feedback.

  6. Knowl 6 — Harness scaling revises the harness for the current task

    algorithm

    Harness scaling spends additional inference compute revising the harness for a single evaluation task, rather than searching for one harness shared across a task set. It first runs the agent on the task using an initial harness. A meta agent then uses the task, the previous harness, and the previous trajectory to create a revised harness, under which the agent runs again. When unit-test feedback is available, the previous rollout’s outcome is also provided to the meta agent, and a successful trajectory can be selected; without tests, the latest trajectory is returned. The process treats task-specific harness adaptation as a harness-level counterpart to test-time scaling, not as evidence that a revision will transfer to other tasks.

  7. Knowl 7 — Trajectory-search baselines use the same budget in different ways

    model/method

    The paper compares two fixed-harness test-time baselines. Parallel sampling generates independent trajectories for the same task and, without unit tests, uses a model-based judge to select one; with unit tests, a successful trajectory can be selected using the verifier. Sequential refinement instead conditions each attempt on the preceding trajectory, allowing the agent to revise its previous solution. Without unit tests it returns the final attempt; with tests, it exposes the previous attempt’s outcome during refinement and can select a successful trajectory. In the experiments, the shared compute budget was five attempts. These baselines spend extra computation directly on finding a successful task trajectory rather than on changing the harness.

  8. Knowl 8 — The comparison uses a common minimal harness and controlled rollout budget

    experimental setup

    Experiments used Terminal-Bench 2.1, a suite of 89 terminal tasks, with Claude Opus 4.6, GPT-5.4, and GPT-5.4 mini. Unless otherwise specified, models used high reasoning effort and a maximum generation budget of 128,000 tokens; reported results average two independent runs. Every method began with the same minimal harness: a single bash tool, with no skills, middleware, or persistent memory. The harness was also held fixed for the trajectory-search baselines. Methods used a compute budget of five, and harness evolution used one rollout per task per harness iteration. The unit-test-feedback and held-out generalization comparisons report only Claude Opus 4.6 and GPT-5.4.

  9. Knowl 9 — Harness revisions often preserve task-specific fixes rather than general strategies

    empirical result

    The authors’ qualitative inspection found that shared-harness evolution commonly began with prompt or long-term-memory advice aimed at recurring failures, then sometimes added runtime enforcement in middleware, such as turn-budget reminders, tool-output truncation, or finalization checks. Harness scaling more often inserted facts and procedures tied to a particular failure or task, including known bugs, file paths, command sequences, and verification steps. Revisions also targeted inefficient setup and polling, fragile state handling, dependency or data-format mistakes, and verification gaps. These changes sometimes helped when an agent already had a viable solution but wasted time or failed during setup. The authors’ interpretation is that many edits memorize fixes the agent might rediscover during a rollout, while persistent prompt growth can consume context and leave difficult failures unchanged.

  10. Knowl 10 — The benchmark may limit what harness evolution can improve

    limitation

    The paper’s proposed explanations for the modest gains are qualified rather than established results. Terminal-Bench agents already solve many tasks, so remaining failures may reflect model reasoning limits rather than fixable harness weaknesses. The benchmark may also be relatively insensitive to harness design: a shell tool and basic prompt can suffice for many solvable tasks, leaving little performance headroom for changes to prompts, tools, or middleware. The authors therefore argue that harness evolution should also be evaluated on tasks that are difficult enough to leave substantial headroom and whose success depends strongly on specialized tools, skills, or workflows.

Coverage note — The paper’s detailed individual case-study edits are not enumerated separately because they illustrate the broader finding that revisions often encode task-specific fixes rather than transferable strategies.

References

  1. 1.Lakshya A Agrawal, Shangyin Tan, Dilara Soylu, Noah Ziems, Rishi Khare, Krista Opsahl-Ong, Arnav Singhvi, Herumb Shandilya, Michael J Ryan, Meng Jiang, et al. Gepa: Reflective prompt evolution can outperform reinforcement learning. arXiv preprint arXiv:2507.19457, 2025.
  2. 2.Bradley Brown, Jordan Juravsky, Ryan Ehrlich, Ronald Clark, Quoc V Le, Christopher Ré, and Azalia Mirhoseini. Large language monkeys: Scaling inference compute with repeated sampling. arXiv preprint arXiv:2407.21787, 2024.
  3. 3.Omar Khattab, Arnav Singhvi, Paridhi Maheshwari, Zhiyuan Zhang, Keshav Santhanam, Sri Vardhamanan, Saiful Haq, Ashutosh Sharma, Thomas T Joshi, Hanna Moazam, et al. Dspy: Compiling declarative language model calls into self-improving pipelines. arXiv preprint arXiv:2310.03714, 2023.
  4. 4.Yoonho Lee, Roshen Nair, Qizheng Zhang, Kangwook Lee, Omar Khattab, and Chelsea Finn. Meta-harness: End-to-end optimization of model harnesses, March 2026. URL http://arxiv.org/abs/2603.28052.
  5. 5.Xiaochuan Li, Ryan Ming, Pranav Setlur, Abhijay Paladugu, Andy Tang, Hao Kang, Shuai Shao, Rong Jin, and Chenyan Xiong. Benchmark test-time scaling of general llm agents. arXiv preprint arXiv:2602.18998, 2026.
  6. 6.Jiahang Lin, Shichun Liu, Chengjun Pan, Lizhi Lin, Shihan Dou, Zhiheng Xi, Xuanjing Huang, Hang Yan, Zhenhua Han, Tao Gui, et al. Agentic harness engineering: Observability-driven automatic evolution of coding-agent harnesses. arXiv preprint arXiv:2604.25850, 2026.
  7. 7.Ryan Lopopolo. Harness engineering: Leveraging codex in an agent-first world, February 2026. URL https://openai.com/zh-Hans-CN/index/harness-engineering/.
  8. 8.Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, et al. Self-refine: Iterative refinement with self-feedback. Advances in Neural Information Processing Systems, 36:46534–46594, 2023.
  9. 9.Mike A Merrill, Alexander G Shaw, Nicholas Carlini, Boxuan Li, Harsh Raj, Ivan Bercovich, Lin Shi, Jeong Yeon Shin, Thomas Walshe, E Kelly Buchanan, et al. Terminal-bench: Benchmarking agents on hard, realistic tasks in command line interfaces. arXiv preprint arXiv:2601.11868, 2026.
  10. 10.Alexander Novikov, Ngân Vu, Marvin Eisenberger, Emilien Dupont, Po-Sen Huang, Adam Zsolt Wagner, Sergey Shirobokov, Borislav Kozlovskii, Francisco JR Ruiz, Abbas Mehrabian, et al. Alphaevolve: A coding agent for scientific and algorithmic discovery. arXiv preprint arXiv:2506.13131, 2025.
  11. 11.Krista Opsahl-Ong, Michael J Ryan, Josh Purtell, David Broman, Christopher Potts, Matei Zaharia, and Omar Khattab. Optimizing instructions and demonstrations for multi-stage language model programs. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 9340–9366, 2024.
  12. 12.Charlie Snell, Jaehoon Lee, Kelvin Xu, and Aviral Kumar. Scaling llm test-time compute optimally can be more effective than scaling model parameters. arXiv preprint arXiv:2408.03314, 2024.
  13. 13.Vivek Trivedy. Improving deep agents with harness engineering, February 2026. URL https://www.langchain.com/blog/improving-deep-agents-with-harness-engineering.
  14. 14.Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. Self-consistency improves chain of thought reasoning in language models. arXiv preprint arXiv:2203.11171, 2022.
  15. 15.John Yang, Carlos Jimenez, Alexander Wettig, Kilian Lieret, Shunyu Yao, Karthik Narasimhan, and Ofir Press. Swe-agent: Agent-computer interfaces enable automated software engineering. Advances in Neural Information Processing Systems, 37:50528–50652, 2024.
  16. 16.Mert Yuksekgonul, Federico Bianchi, Joseph Boen, Sheng Liu, Pan Lu, Zhi Huang, Carlos Guestrin, and James Zou. Optimizing generative ai by backpropagating language model feedback. Nature, 639(8055):609–616, 2025.
  17. 17.Jiayi Zhang, Yongfeng Gu, Jianhao Ruan, Maojia Song, Yiran Peng, Zhiguang Han, Jinyu Xiang, Zhitao Wang, Caiyin Yang, Yixi Ouyang, et al. Harnessing agentic evolution. arXiv preprint arXiv:2605.13821, 2026.
  18. 18.Qizheng Zhang, Changran Hu, Shubhangi Upasani, Boyuan Ma, Fenglu Hong, Vamsidhar Kamanuru, Jay Rainton, Chen Wu, Mengmeng Ji, Hanchen Li, et al. Agentic context engineering: Evolving contexts for self-improving language models. arXiv preprint arXiv:2510.04618, 2025.
  19. 19.Andrew Zhao, Daniel Huang, Quentin Xu, Matthieu Lin, Yong-Jin Liu, and Gao Huang. Expel: Llm agents are experiential learners. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 38, pages 19632–19642, 2024.

Citation

MLA
Wang, Y., et al. “Rethinking the Evaluation of Harness Evolution for Agents”. arXiv, 2026, http://arxiv.org/abs/2607.12227v2.
APA
Wang, Y., Zhu, H., Hu, Z., Yuan, Y., Chen, Z., Senthil, S., Hajishirzi, H., Tsvetkov, Y., Dasigi, P., & Xiao, T. (2026). Rethinking the Evaluation of Harness Evolution for Agents. arXiv. http://arxiv.org/abs/2607.12227v2
Chicago
Wang, Y., H. Zhu, Z. Hu, et al. 2026. “Rethinking the Evaluation of Harness Evolution for Agents”. arXiv. http://arxiv.org/abs/2607.12227v2.
Harvard
Wang, Y. et al. (2026) “Rethinking the Evaluation of Harness Evolution for Agents”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2607.12227v2.
Vancouver
1. Wang Y, Zhu H, Hu Z, Yuan Y, Chen Z, Senthil S, Hajishirzi H, Tsvetkov Y, Dasigi P, Xiao T (2026) Rethinking the Evaluation of Harness Evolution for Agents. arXiv

BibTeX

@article{wang2026rethinking,
  title = {Rethinking the Evaluation of Harness Evolution for Agents},
  author = {Wang, Yike and Zhu, Huaisheng and Hu, Zhengyu and Yuan, Yige and Chen, Zhengyu and Senthil, Shakti and Hajishirzi, Hannaneh and Tsvetkov, Yulia and Dasigi, Pradeep and Xiao, Teng},
  year = {2026},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2607.12227v2},
  eprint = {2607.12227}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/