RRSI: Regularized Recursive Self-Improvement of Agent Harnesses
Peng Xia1,2,∗^{1,2,*}1,2,∗, Rujun Han1^{1}1, Zifeng Wang1^{1}1, Yanfei Chen1^{1}1, Yufan Zhang1^{1}1, Yoonho Lee3^{3}3, Chengsong Huang4^{4}4, Han Yu1^{1}1, Zhongying CuiZhu1^{1}1, Yifei Ming1^{1}1, Huaxiu Yao2^{2}2, Burak Gokturk1^{1}1, Tomas Pfister1^{1}1, Chen-Yu Lee1^{1}1
1^{1}1 Google Research Cloud AI Research
2^{2}2 UNC-Chapel Hill
3^{3}3 Stanford University
4^{4}4 Washington University in St. Louis
2^{2}2 UNC-Chapel Hill
3^{3}3 Stanford University
4^{4}4 Washington University in St. Louis
∗^{*}∗ This work was done while Peng was a Student Researcher at Google Cloud AI Research.
Abstract
An LLM agent’s capability is largely magnified by its harness, namely the prompts, control flow, tooling, memory, and context management surrounding the frozen backbone model.
Recent methods increasingly automate this process by iteratively proposing and selecting component-wise edits of an agent harness, practically establishing a form of recursive self-improvement (RSI) at the agent-system level. However, such recursive evolution may overfit by memorizing the training tasks, showing large in-distribution gains that shrink or even vanish on out-of-distribution benchmarks.
We introduce Regularized Recursive Self-Improvement of Agent Harnesses (RRSI{\mathchoice{\text{R{\scriptsize RSI}}}{\text{R{\scriptsize RSI}}}{\text{R{\scriptscriptstyle RSI}}}{\text{RRSI}}}RRSI), which incorporates the principles of regularizations into harness self-improvement by constraining the evolution candidate proposal and selection.
The proposer operates with a temporally annealed budget, limiting how many edits a candidate can bundle, and it encourages unexplored trajectories based on evolution history. The selector is equipped with a critic and a pruner: the critic screens benchmark-specific proposals, while the pruner, removes changes that are too small, too expensive, or no longer useful. Together these constraints favor reusable agent mechanisms over benchmark-specific ones or even noises.
Across eight benchmarks spanning coding, agentic workspace and engineering design tasks, RRSI{\mathchoice{\text{R{\scriptsize RSI}}}{\text{R{\scriptsize RSI}}}{\text{R{\scriptscriptstyle RSI}}}{\text{RRSI}}}RRSI gains up to 14.1 points on the split it evolves against and up to 4.7 points on the five out-of-distribution benchmarks, while producing a harness that runs on 30% fewer policy tokens than the unregularized evolution.
github.com/google-research/rrsi \textcolor{#C6C6C6}{\rule[0.02em]{0.6pt}{0.9em}}
regularized-rsi.com1. Introduction
Modern LLM agents are systems rather than standalone models ([1, 2]). A frozen backbone model is wrapped in a harness of prompts, control flow, tool interfaces, memory and context management. Agent harness decides whether the same model reads the right file before editing it, recovers from a failed command, manages efficient working context, and writes its findings into the deliverables. Much recent progress in agent products came from harness engineering rather than from new model weights ([3, 4, 5]). However, this engineering relies on manual efforts, where humans inspect failed trajectories and tweak the scaffold by hand, so progress is limited by how many trajectories an engineer can read.
Recent methods automate this loop by using LLMs to optimize harness components from task feedback ([6, 7, 8, 9, 10, 11, 12, 13, 14, 15]). Such iterative harness evolution provides a practical form of recursive self-improvement (RSI) ([16, 17, 18, 19]) at the agent-system level, where feedback from the current system is used to improve the harness that shapes its subsequent behavior. However, as illustrated in Figure 1 (a), test-time harness evolution repeatedly proposes and selects edits using feedback from a finite evolve set, creating an adaptive overfitting risk: evolve-set performance may improve without corresponding gains on unseen tasks. Recent studies observe substantial gaps between evolution and held-out performance, and show that apparent improvements can arise from task-specific fitting or increased test-time computation rather than reusable mechanisms ([20, 21, 22]). Accordingly, recent works explicitly separate evolution and evaluation tasks to measure generalization ([23, 24, 25]). We therefore study the generalization problem in recursive self-improvement, which is defined as evolved harness transferring to unseen benchmarks with different task descriptions, tool interfaces, or verifiers.
Our study shows that overfitting can arise through several coupled behaviors ([25, 26]). The evolution search may encode benchmark-specific patterns, promote candidates favored by the evaluation noise, or accumulate complexity that improves evolve-set scores without improving the underlying agent mechanism. These benchmark-specific fitting, noise chasing, and complexity accumulation all widen the evolve-to-transfer gap. Inspired by these observations, our solution regularizes how recursive harness improvements use finite and noisy feedback.
We introduce RRSI{\mathchoice{\text{R{\scriptsize RSI}}}{\text{R{\scriptsize RSI}}}{\text{R{\scriptscriptstyle RSI}}}{\text{RRSI}}}RRSI, a framework for regularizing the RSI of agent harness that keeps the harness fully editable while constraining how finite evolve-set feedback guides the search. As illustrated in Figure 2, RRSI{\mathchoice{\text{R{\scriptsize RSI}}}{\text{R{\scriptsize RSI}}}{\text{R{\scriptscriptstyle RSI}}}{\text{RRSI}}}RRSI regularizes both sides of the evolution loop: it encourages simpler and more reusable edits when proposing candidates, and applies robust selection criteria to avoid retaining improvements driven by benchmark-specific signals, evaluation noise, or unnecessary complexity. In this way, RRSI{\mathchoice{\text{R{\scriptsize RSI}}}{\text{R{\scriptsize RSI}}}{\text{R{\scriptscriptstyle RSI}}}{\text{RRSI}}}RRSI favors edits that transfer beyond evolution set without restricting which harness components may be updated.
We evaluate RRSI{\mathchoice{\text{R{\scriptsize RSI}}}{\text{R{\scriptsize RSI}}}{\text{R{\scriptscriptstyle RSI}}}{\text{RRSI}}}RRSI on eight benchmarks spanning three domains that differ in task type, tooling and verifier. In each domain the harness is evolved on a single suite, and is then run unchanged on held-out benchmarks. As shown in Figure 1 (b–d), it gains up to 14.1 points on the evolving split and improves all six held-out splits, by up to 4.7 points out of distribution, on fewer policy tokens than unregularized evolution spends. More importantly, these gains generalize beyond the environment used for evolution. RRSI{\mathchoice{\text{R{\scriptsize RSI}}}{\text{R{\scriptsize RSI}}}{\text{R{\scriptscriptstyle RSI}}}{\text{RRSI}}}RRSI retains its improvements across substantially different tasks and evaluation settings, indicating that it learns broadly useful harness changes. More importantly, RRSI{\mathchoice{\text{R{\scriptsize RSI}}}{\text{R{\scriptsize RSI}}}{\text{R{\scriptscriptstyle RSI}}}{\text{RRSI}}}RRSI generalizes across held-out environments, outperforming the average prior baseline by up to 22.9%.
Our contributions are threefold: (1) We identify the overfitting as a key challenge in harness-based recursive self-improvement. (2) We propose RRSI{\mathchoice{\text{R{\scriptsize RSI}}}{\text{R{\scriptsize RSI}}}{\text{R{\scriptscriptstyle RSI}}}{\text{RRSI}}}RRSI, which regularizes both proposal and selection during harness evolution while keeping each harness component editable. (3) Across eight benchmarks in three domains, RRSI{\mathchoice{\text{R{\scriptsize RSI}}}{\text{R{\scriptsize RSI}}}{\text{R{\scriptscriptstyle RSI}}}{\text{RRSI}}}RRSI improves both transfer and efficiency, showing the effectiveness of our proposed approach.
2. Preliminaries
Agents and Harnesses. We consider an agent A=(π,H)A = (\pi, H)A=(π,H) built from a backbone policy π\piπ and a harness HHH. The harness is everything around the weights ([1, 2]): the system and task prompts, the control flow that decides when the agent plans, acts, reflects or stops, the tool interfaces and their descriptions, the memory and skill files the agent may consult, and the context management that decides what the policy sees at each step. Given a task xxx with its environment, the agent produces a trajectory τ∼A(⋅∣x)\tau \sim A(\cdot \mid x)τ∼A(⋅∣x) and a deliverable, which a verifier scores as r(x,τ)∈[0,1]r(x, \tau) \in [0, 1]r(x,τ)∈[0,1]. The verifier can be a unit-test suite in coding environments or a LLM-as-a-judge program in agentic workspace environments. For a task set D\mathcal{D}D, we measure task performance and policy-token cost as
where c(τ)c(\tau)c(τ) is the number of policy tokens consumed by the trajectory.
Harness Evolution. Harness evolution treats HHH as the optimization variable while keeping the backbone policy fixed ([8]). Most methods instantiate the same generic loop. At round ttt, the current harness HtH_tHt is executed on an evolve set Devolve\mathcal{D}_{\mathrm{evolve}}Devolve to obtain trajectories; these trajectories are summarized into feedback Ft\mathcal{F}_tFt; a proposer LLM generates candidate harnesses; the candidates are evaluated on the same evolve set; and the best candidate is selected as the next incumbent. Abstractly,
where P0P_0P0 denotes the unconstrained proposal process and S^\hat{S}S^ is the empirical score obtained from a finite number of stochastic agent runs. With kkk trials per task, we use
Unlike ordinary evaluation, this reuse of Devolve\mathcal{D}_{\mathrm{evolve}}Devolve is adaptive: the candidates proposed at round ttt depend on measurements obtained from the same tasks in earlier rounds. Harness evolution can therefore be viewed as adaptive empirical optimization over an unusually expressive search space.
3. RRSI
We study RSI through iterative harness evolution, where feedback from the current agent system is repeatedly used to propose and select modifications to the harness. RRSI{\mathchoice{\text{R{\scriptsize RSI}}}{\text{R{\scriptsize RSI}}}{\text{R{\scriptscriptstyle RSI}}}{\text{RRSI}}}RRSI follows this recursive improvement process and keeps the harness edit space open, but regularizes how the evolution moves through that space. The key idea is to translate regularization principles from machine learning into an adaptive harness search: sparse updates limit how many mechanisms can change in response to one round of feedback, evidence-aware credit assignment prevents the search from repeatedly spending its capacity on hypotheses it has already falsified, and conservative selection prevents leakage, evaluation noise, or unjustified resource growth from becoming permanent harness state.
3.1 A Regularization View of Harness Evolution
Let Ω(H)\Omega(H)Ω(H) denote the set of harnesses reachable from HHH by arbitrary source edits. RRSI{\mathchoice{\text{R{\scriptsize RSI}}}{\text{R{\scriptsize RSI}}}{\text{R{\scriptscriptstyle RSI}}}{\text{RRSI}}}RRSI deliberately leaves Ω(H)\Omega(H)Ω(H) open: prompts, control flow, configuration, context management, tools, skills, memory, and subagents may all be modified, added, or removed. Instead of restricting this hypothesis space directly, we regularize the search trajectory through it. At each round ttt, the proposer uses feedback from the finite evolve set to generate candidate edits to the current harness HtH_tHt, and the selector determines which, if any, should replace the incumbent.
This view separates two complementary forms of regularization. On the proposal side, we constrain how much adaptive capacity can be exercised in a single round and where that capacity is spent. On the selection side, we constrain which empirical improvements are strong enough, efficient enough, and sufficiently free of leakage to survive. Our complexity control framework takes inspiration from three classical regularization approaches ([27, 28, 29]), and we explain the analogy below. The edit budget is the closest to an L0L_0L0-style cardinality constraint as it directly limits the number of independently active edits in an update. Structural pruning is analogous to Lasso/L1L_1L1-style sparsification because persistently unproductive components are removed from the retained harness, producing a sparser structure. Complexity-aware acceptance is analogous to Ridge/L2L_2L2-style shrinkage since it suppresses unchecked growth in the aggregate resource footprint without requiring any particular component to be eliminated. Detailed algorithm description can be found in Appendix C.
3.2 Regularizing the Proposal Distribution
The proposal distribution determines how aggressively the search can respond to feedback from the evolve set. RRSI{\mathchoice{\text{R{\scriptsize RSI}}}{\text{R{\scriptsize RSI}}}{\text{R{\scriptscriptstyle RSI}}}{\text{RRSI}}}RRSI regularizes it in three ways: it anneals how much update capacity a single round may exercise, it makes credit assignment evidence-aware over the whole run, and it structures where that capacity is spent.
L0L_0L0-Style Annealed Update Sparsity.
An unconstrained proposer can bundle many unrelated modifications into one candidate. Such candidates have high effective capacity: they can fit more idiosyncrasies of the current feedback, and any measured change is difficult to attribute to a particular mechanism. We therefore cap the number of independently attributable edits that may be included in one proposal. At round ttt of a TTT-round run, this budget is
The schedule decreases from bmaxb_{\max}bmax to bminb_{\min}bmin: early rounds may combine several coordinated changes to discover new mechanisms, whereas later rounds become increasingly sparse and attributable. This is our most direct classical analogy: if the independently attributable edits in a candidate are represented by binary activity indicators, the budget bounds their cardinality, i.e., an L0L_0L0-style constraint on the update. The analogy applies to update sparsity rather than to a fixed model parameter vector; the edit pool can change across rounds, and we do not optimize an L0L_0L0-penalized objective.
Evidence-Aware Credit Assignment.
Constraining the size of an update only helps if the search knows what earlier updates established. Every evaluation is another adaptive look at the same finite evolve set, so repeatedly testing hypotheses that earlier rounds already falsified spends search capacity without adding useful evidence ([30]). RRSI{\mathchoice{\text{R{\scriptsize RSI}}}{\text{R{\scriptsize RSI}}}{\text{R{\scriptscriptstyle RSI}}}{\text{RRSI}}}RRSI therefore records, for every evaluated candidate, the component it modifies, the hypothesis it tests, the source diff, the resulting score and cost changes, and whether the candidate was accepted. The proposer conditions on this history in later rounds: rejected mechanisms remain negative evidence, while successful mechanisms retain explicit credit. As later rounds allow fewer edits per candidate, it becomes easier to identify which change is responsible for an observed improvement.
Structured Exploration.
The same history reveals when the proposer has collapsed onto a narrow edit family, for example repeatedly rewriting prompts while leaving agent structural mechanisms untouched. We treat the search as stalled when its progress over the previous www rounds remains within the empirical noise band δ\deltaδ. During a stall, a small portion of the proposal budget is reserved for components that have not yet been exercised in the run. This plays a role similar to diversity or entropy regularization: it redirects limited proposal capacity toward underexplored mechanisms without changing which mechanisms the harness is allowed to contain ([31]).
3.3 Regularizing Candidate Selection
Standard harness evolution can promote the candidate with the largest measured score even when that score reflects explicit leakage, stochastic variation, or costly growth. RRSI{\mathchoice{\text{R{\scriptsize RSI}}}{\text{R{\scriptsize RSI}}}{\text{R{\scriptscriptstyle RSI}}}{\text{RRSI}}}RRSI retains the same empirical objective but regularizes which candidates are allowed to become permanent state. A candidate must satisfy several non-compensatory criteria before its score can justify replacing the incumbent.
Leakage Screening.
Before full evaluation, a critic reads each candidate diff and rejects edits that explicitly encode task names, entity names, task-specific values, answers, or other logic specific to the evolve benchmark, as well as edits that add inert machinery. The screen targets benchmark-specific content rather than particular harness components: generic prompt or tool-description improvements remain valid candidates. Screening before evaluation is important because a leaking candidate never receives the inflated evolve-set score that could make it attractive to subsequent rounds.
Stability-Aware Acceptance.
Repeatedly selecting among noisy evaluations can convert stochastic winners into permanent search state. Before evolution, we repeatedly evaluate the unchanged base harness and estimate an empirical noise band δ\deltaδ. Let S⋆S^\starS⋆ denote the best evolve-set score observed so far. A candidate must satisfy the noise-adjusted floor
The floor prevents the search from walking downhill through a sequence of regressions that are individually small enough to be mistaken for noise. More broadly, it makes selection conservative to fluctuations induced by repeated stochastic evaluation on the same evolve set ([30]).
Ridge/L2L_2L2-Style Complexity-Aware Acceptance.
For a candidate H′H'H′ relative to the current harness HtH_tHt, let
For a candidate whose gain exceeds the noise band, ΔS>δ\Delta S>\deltaΔS>δ, we require
Here, β0\beta_0β0 sets the cost increase tolerated for a negligible score gain, while β1\beta_1β1 controls how much additional cost is allowed as the measured improvement increases. We select these values on the evolve set and keep them fixed for all transfer evaluations. Thus additional inference cost must be justified by measurable performance improvement. This process is analogous to Ridge/L2L_2L2-style shrinkage: it discourages unconstrained growth in the overall magnitude of the solution, which is represented by the harness's aggregate resource footprint in our approach. As a shrinkage method, it does not require any particular component to be removed for sparsity. We use policy-token cost as a common measurable proxy for this footprint. This is an analogy to Ridge's non-sparsifying complexity control. The detailed rule for candidates whose measured change falls within the noise band is deferred to the Appendix C.3.
Lasso/L1L_1L1-Style Structural Pruning.
The annealed budget in Equation 4 sparsifies each update; pruning sparsifies the retained harness. RRSI{\mathchoice{\text{R{\scriptsize RSI}}}{\text{R{\scriptsize RSI}}}{\text{R{\scriptscriptstyle RSI}}}{\text{RRSI}}}RRSI tracks whether recently exercised components have produced a strictly positive measured gain over a fixed pruning window. Components that remain unproductive are reported to the proposer as deletion targets in subsequent rounds. This process imitates the Lasso/L1L_1L1-style sparsification: mechanisms with insufficient evidence of utility are removed entirely, so the retained harness becomes structurally sparser rather than merely cheaper in aggregate. The correspondence is again qualitative, i.e., Lasso reduces the number of parameters through L1L_1L1 regularization, whereas our pruning rule deletes discrete harness components based on their observed contribution. The shared intuition is selective sparsification: a mechanism must continue to earn its place rather than persist simply because score-only evolution has no incentive to remove it ([27]).
4. Experiments
We evaluate RRSI{\mathchoice{\text{R{\scriptsize RSI}}}{\text{R{\scriptsize RSI}}}{\text{R{\scriptscriptstyle RSI}}}{\text{RRSI}}}RRSI on eight benchmarks spanning three domains: Terminal-Bench 2.1 and SWE-bench Verified for coding, Harvey LAB, JobBench, GDPval and APEX-Agents for agentic workspace tasks, and EngDesign and Frontier-Eng for engineering design. Our experiments address the following questions: 1) How does RRSI{\mathchoice{\text{R{\scriptsize RSI}}}{\text{R{\scriptsize RSI}}}{\text{R{\scriptscriptstyle RSI}}}{\text{RRSI}}}RRSI compare to state-of-the-art harness evolution methods? 2) Do the gains transfer to in-distribution held-out tasks and to out-of-distribution benchmarks that the search never saw? 3) What is the contribution of different components? 4) How does the harness evolve over a run, and what does it cost in tokens?
4.1 Experimental Setup
Environments. We evolve harnesses in three types of tasks, i.e., coding tasks, agentic workspace tasks and engineering design tasks. For coding, Terminal-Bench 2.1 ([32]) is a suite of 89 containerized terminal tasks in which the agent drives a real shell and is verified by the task's own unit tests. For agentic workspace tasks, Harvey LAB ([33]) is a legal-work benchmark spanning 25 practice areas. It is split into a fixed evolve set of 120 tasks and a pristine in-distribution held-out set of 40 tasks. For engineering design, EngDesign ([34]) contributes 61 design tasks, each graded by its own frozen simulator rather than by a judge model. To test its generalization capability, we additionally evaluate on out-of-distribution (OOD) held-out benchmarks: SWE-bench Verified ([35]) for repository-level bug fixing on the coding task, JobBench ([36]), GDPval ([37]) and APEX-Agents ([38]) on the agentic workspace task, and Frontier-Eng ([39]) on the engineering design task.
Baselines. We compare against the unevolved base harness H0H_0H0 that every run starts from, and against four recent harness evolution methods, Meta-Harness ([8]), AHE ([9]), TTHE ([11]) and HarnessX ([13]). All the baselines start from the same H0H_0H0 and share the frozen policy, the evolve set and the candidate budget. The detailed descriptions of baselines are given in Appendix B.
Implementation Details. The policy is frozen throughout Claude Opus 4.8 ([40]) across all three domains. The proposer, the analyst that writes the cross-round failure feedback and the leakage critic are all Claude Opus 4.8. The base harness we used are Terminus-2 (for coding) ([32]), a ReAct loop ([41]) over an MCP tool gateway, a dynamic toolbelt ([38]), and ReSum-style context management (for Harvey LAB and EngDesign). Further hyperparameters are given in Appendix D.1.
4.2 Main Results
RRSI{\mathchoice{\text{R{\scriptsize RSI}}}{\text{R{\scriptsize RSI}}}{\text{R{\scriptscriptstyle RSI}}}{\text{RRSI}}}RRSI improves every split outside the evolve set, in all three domains. Figure 3 reports every number against the unevolved base harness H0H_0H0 measured in the same window, so no gain can be attributed to drift in the evaluation infrastructure. The evolve-set gains are 6.0 points on Terminal-Bench 2.1, 4.9 on EngDesign and 1.1 on Harvey LAB. What matters is what remains once the harness leaves those splits. SWE-bench Verified gains 1.8 points although repository-level bug fixing was never scored. The in-distribution held-out split of Harvey LAB gains 2.3, and the three out-of-distribution agentic benchmarks gain between 3.5 and 4.7 points, 7.2% to 13.1%. Frontier-Eng gains 4.3 Medal points, a 24.3% relative improvement. No held-out split regresses anywhere, which is the failure a memorizing harness produces.
RRSI{\mathchoice{\text{R{\scriptsize RSI}}}{\text{R{\scriptsize RSI}}}{\text{R{\scriptscriptstyle RSI}}}{\text{RRSI}}}RRSI consistently outperforms baselines on all held-out datasets. Table 1 runs the four prior methods from the same H0H_0H0 on the same evolve split under the same candidate budget. Every one of them works well on evolve set. The performance on in-distribution held-out split is quite similar. The separation appears out of distribution, and there the ranking inverts. Meta-Harness, the strongest baseline on the evolve split, adds 0.9 points to the out-of-distribution average; HarnessX lands on the base one; AHE and TTHE finish below the harness they started from, TTHE by 1.7 points. RRSI{\mathchoice{\text{R{\scriptsize RSI}}}{\text{R{\scriptsize RSI}}}{\text{R{\scriptscriptstyle RSI}}}{\text{RRSI}}}RRSI posts the smallest evolve-set gain of any evolved harness and the only out-of-distribution average that clears H0H_0H0 by more than a point, 43.6 against 39.7, which is the trade the regularizers are designed to make.
The transfer is not an artifact of judge-mediated grading or of a shared task format. Harvey LAB, JobBench and GDPval are all scored by a judging model, so a harness could in principle raise its score by writing the way a judge rewards rather than by producing better work. The engineering design instance closes that route: each EngDesign and Frontier-Eng task is graded by its own simulation or testbench, the grading is deterministic, and a design either meets the stated constraints or does not. The gains survive there unchanged, and deterministic grading also removes judge variance from the measurement.
4.3 Analysis
The main results establish that the evolved harnesses transfer; this section asks what produced that property, whether it depends on the backbone the search was run with, and what it costs. Unless stated otherwise, every run below uses the agentic workspace instance and shares the base harness, policy, evolve split, round count and candidate budget of the main experiment, so that arms differ only in the factor under study.
Ablation Analysis. We ablate the two groups of regularizers, (i) the proposal-side constraints and (ii) the acceptance-side constraints. As shown in Table 2, removing either group raises the evolve-set score and lowers transfer. Without the acceptance constraints the evolve-set score rises from 90.5 to 91.5 while the out-of-distribution average falls from 43.6 to 41.0 and token cost rises by half, showing that an unconstrained selection rule spends most of its accepted edits on noise and on context rather than on mechanism. Removing the proposal constraints costs only 0.2 points on the evolve split but 1.7 out of distribution, suggesting that steering where the search looks matters even when nothing is rejected. The most significant degradation comes from removing both, which lifts the evolve-set score to 92.8, the highest of any arm, and leaves the out-of-distribution average at 40.3, within a point of the unevolved harness, at 3.80 million tokens per trial against our 2.42.
RRSI{\mathchoice{\text{R{\scriptsize RSI}}}{\text{R{\scriptsize RSI}}}{\text{R{\scriptscriptstyle RSI}}}{\text{RRSI}}}RRSI is not tied to one policy family. To test whether the gains from regularized harness evolution depend on the policy used during search, we independently run the coding evolution with two policy models from different families: Claude Opus 4.8 and Gemini 3.5 Flash ([42]). For each policy, we start from the same coding harness, evolve only on Terminal-Bench 2.1, and evaluate the resulting harness on both the evolve benchmark and SWE-bench Verified. As shown in Table 3, under Gemini 3.5 Flash, RRSI{\mathchoice{\text{R{\scriptsize RSI}}}{\text{R{\scriptsize RSI}}}{\text{R{\scriptscriptstyle RSI}}}{\text{RRSI}}}RRSI improves Terminal-Bench 2.1 from 64.6 to 78.7 and transfers a 2.2-point gain to SWE-bench Verified. Under Claude Opus 4.8 the pattern is the same: Terminal-Bench 2.1 rises from 74.2 to 80.2 and SWE-bench Verified from 82.0 to 83.8, although the stronger policy starts closer to the ceiling of both suites and leaves less room to gain. In both cases the harness improves the unseen benchmark without ever being scored on it, which suggests that the benefits of RRSI{\mathchoice{\text{R{\scriptsize RSI}}}{\text{R{\scriptsize RSI}}}{\text{R{\scriptscriptstyle RSI}}}{\text{RRSI}}}RRSI are not specific to a particular backbone.
The evolved harness still helps under a backbone the search never used. A harness is a program, not a set of weights, so a mechanism that helps only the policy it was searched against is an artifact of that policy rather than a reusable one. We take the final harness of the coding run, evolved with Gemini 3.5 Flash, and evaluate it unchanged with Gemini 3.1 Flash Lite, a smaller model that never took part in the search. As shown in Table 4, Terminal-Bench 2.1 accuracy rises from 11.2 to 14.6, a 30.4% relative gain against a base score less than a fifth of the search policy's. The mechanisms therefore do not depend on the capability level they were searched at, although the absolute gain is smaller because a weaker backbone leaves fewer tasks within reach of any harness.
RRSI{\mathchoice{\text{R{\scriptsize RSI}}}{\text{R{\scriptsize RSI}}}{\text{R{\scriptscriptstyle RSI}}}{\text{RRSI}}}RRSI produces the lightest harness of any evolved harness. Two regularizers act directly on cost: the L1L_1L1-style budget refuses growth that is not paid for when it is proposed, and the pruning rule removes growth that has stopped being paid for since. No prior method carries either constraint, and Figure 4 (a) shows the consequence: all four sit in the region RRSI{\mathchoice{\text{R{\scriptsize RSI}}}{\text{R{\scriptsize RSI}}}{\text{R{\scriptscriptstyle RSI}}}{\text{RRSI}}}RRSI dominates, spending more policy tokens per trial for a lower out-of-distribution average. AHE is the extreme case, at 3.82 million tokens per trial, 58% more than ours, for 4.4 points less out of distribution. The ordering carries over to trajectory length in Figure 4 (b), where RRSI{\mathchoice{\text{R{\scriptsize RSI}}}{\text{R{\scriptsize RSI}}}{\text{R{\scriptscriptstyle RSI}}}{\text{RRSI}}}RRSI runs 26.3 steps per trial against 27.3 to 34.6 for the prior methods. No evolved harness is as cheap as H0H_0H0, at 1.56 million tokens and 21.2 steps, so evolution does buy part of its gain with test-time compute; the budget decides how much.
5. Related Work
Agent harnesses. The harness, and not only the backbone model, determines what an agent can accomplish: engineering reports from frontier labs describe how prompt structure, tool interfaces, context compaction and recovery logic decide whether a long-running agent finishes a task at all ([1, 2]), and recent analyses argue that harnesses compose and generalize in their own right ([4, 3, 43]). A well-designed harness can even substitute for scale, recovering much of a larger backbone's capability at a fraction of the cost ([26]). This engineering is overwhelmingly manual, and because the best harness is tied to a specific backbone, its cost is paid again with every model release ([44]).
Harness evolution. The closest line of work automates that loop: an LLM proposer rewrites the harness and edits are kept if they raise a benchmark score ([8, 9, 11, 13, 10, 12, 14, 45]), or a single component is evolved, such as skills ([46, 47]), memory ([48, 49, 50, 51]) or a preference signal over rollouts ([52]). This inherits both the mechanisms and the risks of self-improving agents that search over their own code under an empirical fitness signal ([17, 16, 53, 54, 55, 56, 57]). Throughout, the search is driven by the score on the suite it optimizes against, with no term for generalization, and the cost is not hypothetical: reported gains often do not survive a change of suite ([20, 23]), and delta attribution separates edits that install a reusable mechanism from those that merely fit the evolution tasks ([21]). Concurrent work also targets generalization directly, either as an explicit objective of the search ([25]) or by replacing greedy selection with a diversity-preserving archive over candidate harnesses ([58]). Our contribution is orthogonal to what these methods edit. We keep the same open edit space and instead regularize the search dynamics: credit assigned over the full evolution history, task-specific logic filtered before scoring, and acceptance against a noise-adjusted baseline, so what survives is a mechanism rather than a fit to the evolution suite.
6. Conclusion
We study iterative harness evolution as a practical form of recursive self-improvement at the agent-system level, and show that this recursive process itself requires regularization. Because a finite evolve set is reused adaptively across rounds, apparent self-improvement can reflect benchmark-specific fitting, evaluation noise, or unnecessary complexity rather than transferable progress. RRSI{\mathchoice{\text{R{\scriptsize RSI}}}{\text{R{\scriptsize RSI}}}{\text{R{\scriptscriptstyle RSI}}}{\text{RRSI}}}RRSI addresses this problem by regularizing both proposal and selection while leaving the harness edit space open. Across coding, agentic workspace, and engineering design tasks, the resulting harnesses improve held-out and cross-benchmark performance while using less inference cost than unregularized evolution. These results suggest that making agent systems increasingly capable through recursive self-improvement requires controlling not only what can change, but also how repeated feedback is converted into persistent changes.
Limitations
Our study focuses on harness-level recursive self-improvement with frozen backbone models, and therefore does not address settings where model weights are updated during evolution. In addition, RRSI{\mathchoice{\text{R{\scriptsize RSI}}}{\text{R{\scriptsize RSI}}}{\text{R{\scriptscriptstyle RSI}}}{\text{RRSI}}}RRSI still relies on a finite evolve set and several regularization hyperparameters, so its effectiveness may depend on the quality of the feedback signal and the chosen search budget. Finally, although we evaluate transfer across multiple domains, benchmarks, and policy models, broader validation is needed to determine how well the method generalizes to substantially different agent architectures, tool ecosystems, and longer-running self-improvement processes.
Appendix
Contents of Appendix ---
A. Evaluation
This shows how each environment is run and scored. A harness and its baseline are always evaluated in the same window, with the same tool environment, the same judge and the same number of trials.
A.1 Terminal-Bench 2.1
Each task is a container image with a task description, a working directory and a set of unit tests that are hidden from the agent ([32]). The agent drives a real shell through the harness, and a task counts as solved only if the task's own test suite passes after the agent stops, so the reward is exact and cannot be produced by a plausible-looking answer. The reported accuracy is the fraction of the 89 tasks solved in this way. Containers are torn down and rebuilt between arms so that no state carries from one evaluation to the next.
A.2 SWE-bench Verified
Each instance is a real GitHub issue paired with the repository snapshot at the time of the report ([35]). The agent must produce a patch, which is then applied to the snapshot and checked against the instance's fail-to-pass tests, which must go from failing to passing, and its pass-to-pass tests, which must remain passing. The reported resolve rate is the fraction of instances that satisfy both conditions.
A.3 Harvey LAB
Each task provides a folder of source documents in Word, Excel and PDF form and requires the agent to produce deliverable files under exact requested filenames ([33]), which are graded by a strict per-criterion rubric of 20 to 100 independently judged criteria per task, roughly 14,000 criterion verdicts per full evaluation. A criterion is judged in isolation by an LLM judge (Gemini-3.5-Flash ([42])) that reads the produced deliverable together with that single criterion, and the score of a run is the fraction of criteria passed over all tasks, so a task with a long rubric contributes proportionally more evidence than a short one and a missing deliverable fails every criterion it was supposed to satisfy rather than being dropped. The 160 tasks are partitioned once into a 120-task evolve set and a 40-task held-out set, and the partition is fixed for the experiment.
A.4 JobBench
Tasks are drawn from real professional workflows ([36]), each shipping a task folder of input files and a wrapper prompt, with the reference material the agent would need to look up deliberately withheld so that part of the work is genuine retrieval. The harness exposes a filesystem, a code execution tool for producing office and PDF deliverables, and a grounded web search tool. Deliverables are graded by the benchmark's own weighted rubric, and the reported number is the weighted rubric score over the evaluated split. We use an LLM judge (average score of Gemini-3.5-Flash and Claude Opus 4.8).
A.5 GDPval
For each task the deliverable produced by the harness is placed side by side with the human expert deliverable shipped with the benchmark ([37]), a panel of three judges of different provenance picks the better of the two, and the reported number is the win rate against the expert over 185 tasks. The panel combines an open-weight model served locally (Qwen3.6-35B-A3B ([59])) with two proprietary models from different vendors (Claude Sonnet 4.6 ([60]) and Gemini-3.1 Pro ([61])), each pair is judged in both presentation orders to remove position bias, and the verdict for a task is the majority vote of the three. Each judge therefore issues 204 comparisons per harness, and a win rate above 50% means the harness produces the preferred deliverable more often than the human expert it is compared against.
A.6 APEX-Agents
Each task places the agent in a sandboxed world with its own MCP tool surface, covering a filesystem, PDF reading, spreadsheets, mail, chat, calendar, documents and code execution, and spanning three professional domains ([38]). A task is graded by a per-task rubric judged by an LLM judge (Gemini-3.5-Flash), and a task counts as a success under pass@1 only when its rubric is satisfied on the single sampled rollout. We evaluate the full set of 480 tasks and always report over that full denominator, so a task whose rollout is missing because of an infrastructure failure counts as a failure rather than being excluded, which prevents a harness that crashes on hard worlds from looking better than one that attempts them.
A.7 EngDesign
We used the license-free subset of EngDesign ([34]), of which we take the 61 tasks that run without proprietary simulators. Each task states a design goal together with the physical constraints the design must satisfy, and each is graded by its own frozen simulation or testbench rather than by a judge model, so grading is deterministic and every point of variance we measure comes from the policy. Evolution runs on all 61 tasks with no in-distribution held-out split, since the suite is too small to spend tasks on one.
A.8 Frontier-Eng
Frontier-Eng ([39]) collects real-world engineering optimization problems from 26 domains. Each task asks the agent to produce a design or a program that is scored by a frozen task-specific simulator or evaluator on a continuous objective, so as with EngDesign no judge model is involved and grading is deterministic. Because the objectives are not commensurable across tasks, the benchmark reports a Medal Score: for each task the three best feasible results of the frozen v1 snapshot are the gold, silver and bronze thresholds, a submission earns 1, 0.67 or 0.33 for reaching each, and the score is the mean credit over the 47 tasks of the v1 set, which we report as a percentage. We use Frontier-Eng only as an out-of-distribution test surface. Its EngDesign domain reuses tasks from our evolve set and is excluded, and tasks whose evaluation environment could not be built in our sandbox receive no credit in either arm, so 38 of the 47 tasks contribute credit and both arms are scored on exactly the same tasks.
B. Baseline Methods
We briefly summarize the four harness-evolution baselines used in our experiments.
Meta-Harness ([8]). Meta-Harness formulates harness engineering as an outer-loop optimization problem over executable harness code. Its agentic proposer has access to the source code, evaluation scores, and execution traces of previous candidates, and uses this accumulated experience to propose improved harnesses.
Agentic Harness Engineering (AHE) ([9]). AHE uses an observability-driven evolution loop for coding-agent harnesses. It organizes harness components, execution experience, and edit outcomes into explicit representations so that an evolving agent can diagnose failures, propose changes, and evaluate the effects of previous edits.
Test-Time Harness Evolution (TTHE) ([11]). TTHE evolves executable harnesses during test-time adaptation while keeping the underlying model weights fixed. It maintains multiple candidate harnesses, proposes modifications from execution traces, and uses an agentic judge to select a harness that persists to subsequent inputs.
HarnessX ([13]). HarnessX represents an agent harness as a composition of modular, typed primitives spanning components such as prompts, tools, memory, and control flow. Its trace-driven adaptation mechanism uses execution feedback to modify and select harness configurations, enabling the runtime scaffold to evolve over time.
C. Method Details
This section gives the round-level formulation and implementation details omitted from Section 3. It specifies the same proposal- and selection-side regularizers used in the experiments. As in the main text, the L0L_0L0, Lasso/L1L_1L1, and Ridge/L2L_2L2 terminology is used only to indicate analogous roles in complexity control. The procedure does not optimize the corresponding norm-penalized objectives, and heterogeneous harness components are not treated as coordinates of a shared continuous parameter vector.
C.1 Round-Level Formulation
Let Ω(H)\Omega(H)Ω(H) denote the set of harnesses reachable from HHH by arbitrary source edits. RRSI{\mathchoice{\text{R{\scriptsize RSI}}}{\text{R{\scriptsize RSI}}}{\text{R{\scriptscriptstyle RSI}}}{\text{RRSI}}}RRSI leaves Ω(H)\Omega(H)Ω(H) open and instead regularizes the transition through this space. A round takes the form
with Ht+1=HtH_{t+1}=H_tHt+1=Ht if no candidate is admissible. Here Ft\mathcal{F}_tFt is feedback from the current round, Lt\mathcal{L}_tLt is the edit history, btb_tbt is the annealed edit budget from Equation 4, Et\mathcal{E}_tEt contains exploration directives, Bt\mathcal{B}_tBt contains structural pruning targets inferred from recent history, and At\mathcal{A}_tAt is the set of candidates allowed to replace the incumbent.
A run applies Algorithm 1 and Algorithm 2 for t=0,…,T−1t=0,\ldots,T-1t=0,…,T−1, starting from H0H_0H0 with S⋆=S^(H0)S^\star=\hat{S}(H_0)S⋆=S^(H0). Before evolution, the unchanged base harness is evaluated repeatedly to estimate the empirical noise tolerance δ\deltaδ.
C.2 Proposal-Side Bookkeeping
Atomic edit representation.
At round ttt, the proposer drafts a pool EtE_tEt of atomic edits to HtH_tHt, and a candidate applies a subset of that pool. Write this subset as zt∈{0,1}∣Et∣z_t\in\{0,1\}^{|E_t|}zt∈{0,1}∣Et∣, with zt,j=1z_{t,j}=1zt,j=1 when edit jjj is included. The pool is redrawn each round from the open space Ω(Ht)\Omega(H_t)Ω(Ht), so ∣Et∣|E_t|∣Et∣ need not be fixed across rounds. The annealed budget in Equation 4 imposes
Thus btb_tbt limits the number of independently attributable edits bundled into one candidate rather than the set of components that may eventually be modified. This is the most direct of our classical analogies: it is a cardinality constraint on the update, not an L0L_0L0 penalty on a fixed model parameter vector.
Edit history and component-level summaries.
Every atomic edit in an evaluated candidate is tagged with a component ℓ\ellℓ, a hypothesis hhh, and the candidate source diff ddd. A candidate containing multiple edits contributes one history record per edit; all edits in that candidate share the same measured ΔS\Delta SΔS, ΔC\Delta CΔC, and round outcome. Ignoring candidates that fail before a valid measurement is obtained, the history before round ttt can be written
where ai=1a_i=1ai=1 iff the candidate carrying edit iii was selected as the winner of its round and therefore entered the accepted evolution path. Candidates that are admissible but lose to a higher-scoring admissible candidate have ai=0a_i=0ai=0.
Two summaries used by the proposer are
Here Tt\mathcal{T}_tTt is the set of components with at least one measured edit, and gt(ℓ)g_t(\ell)gt(ℓ) is the best recent measured gain associated with component ℓ\ellℓ over the pruning window. Because bundled edits inherit the candidate-level measurement, this evidence becomes more attributable as the edit budget anneals toward one.
Structured exploration state.
Let K\mathcal{K}K denote the editable component vocabulary. In the implementation,
The exploration directive is
where σt\sigma_tσt indicates that progress over the previous www rounds has not exceeded the empirical noise tolerance, Ut\mathcal{U}_tUt contains components not yet exercised by a measured edit, and mdraftm_{\mathrm{draft}}mdraft reserves candidate slots for exploratory edits when the search is stalled.
Structural pruning.
The pruning target set is
Thus a component is marked as unproductive when it has been exercised but has produced no strictly positive measured gain in the recent pruning window. The proposer receives Bt\mathcal{B}_tBt together with any previously accepted edits associated with those components and is instructed to remove unproductive machinery in subsequent proposals. This is analogous in role to Lasso/L1L_1L1-style sparsification because the mechanism acts by deleting discrete structure from the retained harness; it is not an L1L_1L1-penalized continuous optimization problem.
C.3 Selection-Side Bookkeeping
Noise-adjusted floor.
The leakage critic is applied before full evaluation. For every candidate that reaches selection, the first non-compensatory performance requirement is the stability floor from Equation 5,
This permits fluctuations within the empirically calibrated tolerance while preventing the search from accumulating a sequence of small regressions.
Novelty used by the within-band rule.
The shaped rule uses novelty only for structural components. Let
and let Nt(ℓ)N_t(\ell)Nt(ℓ) be the number of previously accepted edit records tagged with component ℓ\ellℓ before round ttt. If comp(H′)\mathrm{comp}(H')comp(H′) is the set of component types touched by candidate H′H'H′, the implementation computes
Hence νt(H′)\nu_t(H')νt(H′) counts distinct structural component types touched by the candidate that have never previously appeared in a winning edit. Prompt, control-flow, configuration, output-plumbing, and context-management edits do not receive this novelty bonus.
Acceptance when the gain exceeds the noise tolerance.
For ΔS>δ\Delta S>\deltaΔS>δ, the selector uses the gain-dependent cost condition from Equation 7,
The rule allows more inference cost only when accompanied by a larger measured improvement. This is the part of complexity-aware acceptance that motivates the Ridge/L2L_2L2-style analogy in the main text: it suppresses unchecked growth in aggregate resource footprint without requiring an individual component to be eliminated. The analogy is functional rather than mathematical; the rule is not a squared-norm penalty.
Acceptance when the gain does not exceed the noise tolerance.
For candidates that pass the stability floor but whose measured gain does not exceed the empirical tolerance, ΔS≤δ\Delta S\le\deltaΔS≤δ, the implementation does not use Equation 7. Instead it applies the shaped admissibility condition
Here ws,wc,wn≥0w_s,w_c,w_n\ge0ws,wc,wn≥0 control, respectively, the contribution of the measured score change, relative inference-cost change, and previously unused structural component types. The purpose of this branch is to avoid treating a small score fluctuation as sufficient evidence by itself. Within this region, reducing cost contributes positively through −wcΔC-w_c\Delta C−wcΔC, and trying a structural mechanism that has never previously entered the accepted evolution path contributes through wnνt(H′)w_n\nu_t(H')wnνt(H′). Depending on the evolution instance, a within-band score change may also contribute through wsΔSw_s\Delta SwsΔS.
The coding instance sets ws=0w_s=0ws=0. Consequently, a score increase that remains within δ\deltaδ cannot by itself make a coding candidate admissible; the candidate must instead obtain sufficient credit from lower cost and/or structural novelty. The agentic-workspace and engineering-design instances use positive wsw_sws. All three weights are fixed for an evolution instance and are reported in Table 5. Equation 17 is an implementation-level tie-breaking/admissibility rule inside the uncertainty region; it is not itself identified with an LpL_pLp penalty.
Domain-specific non-compensatory guards.
After the stability and cost checks, the implementation may apply a domain-specific guard g(Ht,H′)∈{0,1}g(H_t,H')\in\{0,1\}g(Ht,H′)∈{0,1}. The coding and agentic-workspace instances use no additional guard, so g=1g=1g=1. The engineering-design instance additionally rejects a candidate if its valid-output rate falls by more than 0.030.030.03 relative to the incumbent or if its no-submission rate rises by more than 0.020.020.02. These guards prevent a gain in the primary pass-rate objective from compensating for a substantial degradation in basic execution validity.
Final round selection.
A candidate is admissible only if it satisfies the noise-adjusted floor, the appropriate branch of the complexity-aware rule, and all active domain guards. Among admissible candidates, the selector chooses the one with the largest measured score; if none is admissible, the incumbent is retained. The running best score is then updated as S⋆←max(S⋆,S^(Ht+1))S^\star\leftarrow\max(S^\star,\hat{S}(H_{t+1}))S⋆←max(S⋆,S^(Ht+1)).
D. Experiments
D.1 Hyperparameter Setting
RRSI{\mathchoice{\text{R{\scriptsize RSI}}}{\text{R{\scriptsize RSI}}}{\text{R{\scriptscriptstyle RSI}}}{\text{RRSI}}}RRSI introduces a small number of hyperparameters that control update sparsity, exploration, pruning, and the cost–performance trade-off. We select these parameters using only the evolve environment and operational considerations; held-out and OOD benchmarks are not used for tuning. The noise tolerance δ\deltaδ is calibrated from repeated evaluations of the unchanged base harness. The edit-budget parameters (bmin,bmax)(b_{\min},b_{\max})(bmin,bmax) determine how many independent changes can be bundled into one candidate, while www and mdraftm_{\mathrm{draft}}mdraft control when and how strongly the search explores underused components. The pruning window nprunen_{\mathrm{prune}}nprune determines how much recent evidence is required before a component is treated as unproductive. Finally, (β0,β1)(\beta_0,\beta_1)(β0,β1) encode the allowed trade-off between measured gain and additional inference cost. Table 5 lists the values used in each instance. Scores S^\hat{S}S^ are fractions in [0,1][0,1][0,1] and ΔC\Delta CΔC is the relative change in policy tokens per trial, so δ\deltaδ and β1\beta_1β1 are expressed in those units: on the coding instance δ\deltaδ corresponds to 3 passes out of 89×k=17889\times k=17889×k=178 trials, on the agentic workspace instance to 60 criteria out of roughly 14,10014{,}10014,100 criterion verdicts, and on the engineering design instance to 5 passes out of 61×k=24461\times k=24461×k=244 trials. Likewise β1\beta_1β1 corresponds to a 25% token allowance per additional pass (coding), per 100 additional criteria (agentic workspace) and a 10% allowance per additional pass (engineering design).
E. Qualitative Case Study
To complement the aggregate results, we inspect representative decisions made during RRSI{\mathchoice{\text{R{\scriptsize RSI}}}{\text{R{\scriptsize RSI}}}{\text{R{\scriptscriptstyle RSI}}}{\text{RRSI}}}RRSI evolution. Table 6 summarizes several examples from the released trajectories. The complete round-by-round records, including proposals, critic decisions, acceptance decisions, and exact harness diffs, are available on our project website.
These examples provide a more concrete view of the regularization behavior. In particular, the two candidates from the first coding round are superficially similar, yet only the candidate with a sufficiently large measured improvement survives the cost-aware selection rule. Conversely, the round-8 candidate reduces inference cost but is still rejected because its performance falls below the admissible floor. The engineering example shows the complementary case: a small and reusable control-flow correction is retained with little resource growth. Together, these trajectories suggest that RRSI{\mathchoice{\text{R{\scriptsize RSI}}}{\text{R{\scriptsize RSI}}}{\text{R{\scriptscriptstyle RSI}}}{\text{RRSI}}}RRSI does not simply accumulate edits that improve the evolve-set score, but selectively retains changes whose measured benefit is sufficiently robust relative to their complexity.
References
[1] Prithvi Rajasekaran (2026). Harness design for long-running application development. https://www.anthropic.com/engineering/harness-design-long-running-apps.
[2] Ryan Lopopolo (2026). Harness engineering: leveraging Codex in an agent-first world. https://openai.com/index/harness-engineering/.
[3] Weng, Lilian (2026). Harness Engineering for Self-Improvement. lilianweng.github.io. https://lilianweng.github.io/posts/2026-07-04-harness/.
[4] Zhang, Alex and Khattab, Omar (2026). Language model harnesses are compositional generalizers. https://alexzhang13.github.io/blog/2026/harness/.
[5] Karten et al. (2026). Prime agent: A self-improving rlm harness. Prime Intellect Blog.
[6] Lou et al. (2026). Autoharness: improving llm agents by automatically synthesizing a code harness. arXiv preprint arXiv:2603.03329.
[7] Joel Niklaus (2026). Don't Train the Model, Evolve the Harness. https://huggingface.co/spaces/joelniklaus/harness-optimization.
[8] Lee et al. (2026). Meta-harness: End-to-end optimization of model harnesses. The Third Conference on Language Modeling.
[9] Lin et al. (2026). Agentic harness engineering: Observability-driven automatic evolution of coding-agent harnesses. arXiv preprint arXiv:2604.25850.
[10] Zhang et al. (2026). Self-harness: Harnesses that improve themselves. arXiv preprint arXiv:2606.09498.
[11] Nie et al. (2026). TTHE: Test-Time Harness Evolution. arXiv preprint arXiv:2607.08124.
[12] Lee et al. (2026). Recursive Harness Self-Improvement. arXiv preprint arXiv:2607.15524.
[13] Chen et al. (2026). HarnessX: A Composable, Adaptive, and Evolvable Agent Harness Foundry. arXiv preprint arXiv:2606.14249.
[14] Karten et al. (2026). Continual harness: Online adaptation for self-improving foundation agents. arXiv preprint arXiv:2605.09998.
[15] Zhang et al. (2026). DarwinX: Evolving Agent Harnesses Through Natural Selection. arXiv preprint arXiv:2608.07545.
[16] Wang et al. (2025). Huxley-G\\backslash\" odel Machine: Human-Level Coding Agent Development by an Approximation of the Optimal Self-Improving Machine. arXiv preprint arXiv:2510.21614.
[17] Zhang et al. (2026). Darwin Gödel machine: open-ended evolution of self-improving agents. In International Conference on Learning Representations. pp. 104223–104294.
[18] Team et al. (2026). NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness. arXiv preprint arXiv:2609.08183.
[19] RSI-Exam Team (2026). RSI-Exam: Benchmarking Recursive Self-Improvement through Executable Research. https://github.com/aiming-lab/RSI-Exam.
[20] Wang et al. (2026). Rethinking the evaluation of harness evolution for agents. In COLM 2026 The 2nd Workshop on Lifelong Agents: Learning, Aligning, and Evolving.
[21] Ding et al. (2026). What Evolves When We Talk About Harness Evolution?. wenwen-d.github.io. https://wenwen-d.github.io/blog/harness-delta-attribution/.
[22] Lin et al. (2026). Harness updating is not harness benefit: Disentangling evolution capabilities in self-evolving llm agents. arXiv preprint arXiv:2605.30621.
[23] Huang et al. (2026). Evo-Bench: Can Language Models Improve Agent Harness?. arXiv preprint arXiv:2608.09096.
[24] Ke et al. (2026). EVOHARNESSBENCH: Can Your Agents Keep Pace with an Evolving Harness?. arXiv preprint arXiv:2609.04280.
[25] Zhang et al. (2026). HarnessCompass: Guiding Automatic Harness Evolution toward Generalizable and Effective Agent Harnesses. arXiv preprint arXiv:2608.01918.
[26] Yang et al. (2026). Better Harnesses, Smaller Models: Building 90% Cheaper Agents via Automated Harness Adaptation. arXiv preprint arXiv:2607.08938.
[27] Hastie et al. (2009). The elements of statistical learning: data mining, inference, and prediction. Springer.
[28] Goodfellow et al. (2016). Deep learning. MIT press Cambridge.
[29] Louizos et al. (2018). Learning sparse neural networks through L0L_0L0 regularization. In International Conference on Learning Representations.
[30] Dwork et al. (2015). Generalization in adaptive data analysis and holdout reuse. Advances in neural information processing systems. 28.
[31] Haarnoja et al. (2018). Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor. In International conference on machine learning. pp. 1861–1870.
[32] Merrill et al. (2026). Terminal-bench: Benchmarking agents on hard, realistic tasks in command line interfaces. In International Conference on Learning Representations. pp. 40903–40986.
[33] Harvey AI (2026). Harvey LAB: The Legal Agent Benchmark. Announcement: https://www.harvey.ai/blog/introducing-harveys-legal-agent-benchmark. https://github.com/harveyai/harvey-labs/tree/v1.0.
[34] Guo et al. (2025). Toward engineering AGI: Benchmarking the engineering design capabilities of LLMs. Advances in Neural Information Processing Systems.
[35] Jimenez et al. (2024). Swe-bench: Can language models resolve real-world github issues?. In International Conference on Learning Representations. pp. 54107–54157.
[36] Li et al. (2026). JobBench: Aligning Agent Work With Human Will. arXiv preprint arXiv:2605.26329.
[37] Patwardhan et al. (2026). Gdpval: Evaluating ai model performance on real-world economically valuable tasks. In International Conference on Learning Representations. pp. 24005–24040.
[38] Vidgen et al. (2026). APEX-agents. arXiv preprint arXiv:2601.14242.
[39] Chi et al. (2026). Frontier-Eng: Benchmarking self-evolving agents on real-world engineering tasks with generative optimization. arXiv preprint arXiv:2604.12290.
[41] Yao et al. (2022). React: Synergizing reasoning and acting in language models. arXiv preprint arXiv:2210.03629.
[42] Google (2026). Gemini 3.5: frontier intelligence with action. https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-5/.
[43] Wang et al. (2026). Harness Handbook: Making Evolving Agent Harnesses Readable, Navigable, and Editable. arXiv preprint arXiv:2607.13285.
[44] Huang et al. (2026). EnvHarness: Awakening Static Worlds for Agent Learning. arXiv preprint arXiv:2608.19880.
[45] Liu et al. (2026). Adaptive Auto-Harness: Sustained Self-Improvement for Agentic System Deployment on Open-Ended Task Streams. arXiv preprint arXiv:2606.01770.
[46] Yang et al. (2026). Skillopt: Executive strategy for self-evolving agent skills. arXiv preprint arXiv:2605.23904.
[47] Xia et al. (2026). Skillrl: Evolving agents via recursive skill-augmented reinforcement learning. arXiv preprint arXiv:2602.08234.
[48] Tang et al. (2025). Agent kb: Leveraging cross-domain experience for agentic problem solving. arXiv preprint arXiv:2507.06229.
[49] Ouyang et al. (2026). Reasoningbank: Scaling agent self-evolving with reasoning memory. In International Conference on Learning Representations. pp. 94327–94354.
[50] Liu et al. (2026). Evolvemem: Self-evolving memory architecture via autoresearch for llm agents. arXiv preprint arXiv:2605.13941.
[51] Wu et al. (2026). AutoMem: Automated Learning of Memory as a Cognitive Skill. arXiv preprint arXiv:2607.01224.
[52] Pan et al. (2026). Retrospective Harness Optimization: Improving LLM Agents via Self-Preference over Trajectory Rollouts. arXiv preprint arXiv:2606.05922.
[53] Zhang et al. (2026). Hyperagents. arXiv preprint arXiv:2603.19461.
[54] Xia et al. (2026). Agent0: Unleashing self-evolving agents from zero data via tool-integrated reasoning. The Third Conference on Language Modeling.
[55] Xia et al. (2026). MetaClaw: Just Talk–An Agent That Meta-Learns and Evolves in the Wild. arXiv preprint arXiv:2603.17187.
[56] Huang et al. (2026). R-zero: Self-evolving reasoning llm from zero data. In International Conference on Learning Representations. pp. 130770–130790.
[57] Huang et al. (2026). G-Zero: Self-play for open-ended generation from zero data. arXiv preprint arXiv:2605.09959.
[58] Luo et al. (2026). Self-Evolving Agent Harnesses via Gated Semantic Quality-Diversity. arXiv preprint arXiv:2607.13683.
[60] Anthropic (2026). Introducing Claude Sonnet 4.6. https://www.anthropic.com/news/claude-sonnet-4-6.
[61] Google (2026). Gemini 3.1 Pro: Best for complex tasks and bringing creative concepts to life. https://deepmind.google/models/gemini/pro/.







