Rethinking the Evaluation of Harness Evolution for Agents
Yike WangHuaisheng ZhuZhengyu HuYige YuanZhengyu ChenShakti SenthilHannaneh HajishirziYulia TsvetkovPradeep DasigiTeng Xiao
Demonstrates that automated agent scaffolding evolution fails to consistently outperform simple test-time scaling baselines or generalize to held-out tasks when evaluated under equal compute and feedback budgets.
Artificial intelligence agents driven by large language models increasingly rely on external harnesses—the supporting scaffolding of prompts, tools, memory, and control logic that structures how models interact with complex environments. While researchers have proposed automatic harness evolution to iteratively refine this scaffolding using benchmark feedback, existing evaluation protocols often test the evolved harnesses on the very same benchmarks used during optimization. This creates serious risks of task overfitting and obscures whether performance gains stem from genuine improvements in harness architecture or simply from spending extra search and inference compute during testing.
The article evaluates the true efficacy and generalizability of automatic harness evolution for AI agents. It investigates whether automatic harness modifications outperform standard, lightweight test-time scaling baselines under matched computational and feedback budgets.
The authors conducted a controlled empirical evaluation across 89 terminal coding tasks from the Terminal-Bench 2.1 benchmark using leading frontier models, including Claude Opus 4.6, GPT-5.4, and GPT-5.4 mini. They evaluated four distinct approaches under a unified budget framework: parallel sampling, sequential refinement, dataset-wide harness evolution, and instance-specific harness scaling. The experiments tested performance under three conditions: without access to unit test feedback, with access to unit test feedback, and across a strict data split where harnesses were trained, validated, and tested on completely disjoint task sets.
The findings show that automatic harness evolution fails to consistently outperform basic test-time scaling methods and exhibits severe overfitting. When unit test feedback is unavailable, automatic harness evolution averaged a 67.4% success rate, underperforming both direct baseline sampling at 68.2% and simple parallel sampling at 72.3%. When unit test feedback was provided, parallel sampling reached an 86.0% pass rate and sequential refinement achieved a 91.8% success rate across five attempts, substantially surpassing harness evolution's 75.8% and 86.2% respectively. Most critically, when evaluated on held-out tasks, harness evolution yielded virtually no benefit, providing an average gain of just 0.6 percentage points over the unevolved baseline.
These results indicate that automated harness revision primarily memorizes task-specific facts and fixes rather than discovering generalizable design strategies. The performance improvements reported in prior studies largely reflect the benefits of repeated sampling rather than superior scaffolding design. Consequently, allocating computational budgets directly to solution sampling and refinement is a far more reliable, cost-effective, and lower-risk strategy than dedicating resources to automated harness rewriting, which often introduces context bloat without solving fundamental reasoning failures.
Organizations developing AI agents should prioritize straightforward test-time scaling techniques, such as parallel candidate generation and sequential refinement, over complex automatic harness evolution. Research teams evaluating harness engineering must adopt rigorous protocols that strictly separate training tasks from evaluation tasks and benchmark all gains against compute-matched sampling baselines. Future work should investigate harness design on benchmarks where specialized tooling is strictly required to expand model capabilities.
The conclusions are supported by rigorous multi-model evaluations, but findings are bounded by the specific characteristics of the Terminal-Bench environment, where baseline models already achieve high performance and tasks require only minimal shell tool interfaces. Confidence in the core finding remains high: current automatic harness evolution methods risk severe overfitting and do not reliably deliver generalizable improvements over simpler inference-time search baselines.
- Paper: Automated Design of Agentic Systems, Shengran Hu et al. (2025). Its automated search over agent designs provides a concrete precursor for understanding the harness-evolution methods whose generalizability this paper evaluates.
- Paper: RRSI: Regularized Recursive Self-Improvement of Agent Harnesses, Peng Xia et al. (2026). RRSI directly takes up the overfitting problem identified here, adding regularization and held-out evaluation to make recursive harness improvements generalize.
