Harness-Zero: Harness Distillation via Agent-as-Harness
Haoran Ye1^{1}1, Yuxing Lu2,3^{2,3}2,3, Haonan Dong1^{1}1, Zhaochen Su4^{4}4, Guojie Song1,✉^{1,✉}1,✉
1^{1}1State Key Laboratory of General Artificial Intelligence, School of Intelligence Science and Technology, Peking University
2^{2}2College of Future Technology, Peking University
3^{3}3Google
4^{4}4The Hong Kong University of Science and Technology
1^{1}1State Key Laboratory of General Artificial Intelligence, School of Intelligence Science and Technology, Peking University
2^{2}2College of Future Technology, Peking University
3^{3}3Google
4^{4}4The Hong Kong University of Science and Technology
✉^{✉}✉ Corresponding author
Abstract
Agent harnesses, the external systems that mediate model-environment interaction, can substantially improve agent performance, but their gains remain tied to the harness at deployment. Because the best harness varies across domains, instances, and models, a general-purpose agent must either settle for a suboptimal shared harness or route among an ever-growing set of specialized ones.
We therefore study agent harness distillation: using a domain- or instance-optimized harness as training-time guidance and transferring the behaviors it induces into model weights, so that its gains survive under a single fixed target harness.
The challenge is that the two harnesses differ in action space and available information, so guidance from the optimized harness cannot serve directly as supervision for the target one.
We introduce Harness-Zero, which enables harness distillation through agent-as-harness. Guided by the optimized harness, a harnessing agent corrects student responses before execution in the target harness's action space, turning harness guidance into training demonstrations.
Fine-tuning on the resulting trajectories internalizes harness-induced behavior into the model, so the specialized harness can be removed at deployment.
Our experiments spanning knowledge work, tool use, and science domains show that:
❶ For frontier LLMs using the same evolved harness, agent-as-harness outperforms code-as-harness.
❷ With the specialized harness removed at deployment, Harness-Zero improves the base model's macro-average task success from 23.3% to 44.3%, even exceeding the 41.7% it reaches with that harness still attached.
❸ Harness-Zero recovers harness-induced behaviors absent from the base model, with 82.3% average recovery across 28 patterns in the three domains.
1. Introduction
An LLM agent's capabilities depend on both its model and its harness: the external system that organizes tool use, manages context and state, and controls interaction with the environment ([1, 2]). Harness engineering has become a central lever for improving agent performance. Coding-agent harnesses combine shell access, file systems for persistent memory, subagents, and background jobs ([1, 3]). Research-agent harnesses organize workflows for hypothesis generation, experimentation, and evidence collection ([4]). Context management and experience reuse further support continual learning and long-horizon execution ([5, 6, 7, 8, 9, 10]). Recent methods such as Meta-Harness automate this engineering process by optimizing harness code ([11, 12, 13]).
Harness optimization, however, improves the agent's external scaffolding rather than the model itself, so its gains remain tied to that harness at deployment. Because the best harness varies across domains, instances, and base models ([13, 14, 15]), a general-purpose agent must choose between a shared harness and a collection of specialized ones. A shared harness forgoes some specialized gains ([14, 15, 16]). Maintaining many harnesses instead requires routing and incurs recurring costs in context, model calls, tool calls, and orchestration ([17, 18, 19, 20]). Neither choice moves the discovered harness improvements into the model.
The distillation target in HARNESS-ZERO{\mathchoice{\text{H{\scriptsize ARNESS}-Z{\scriptsize ERO}}}{\text{H{\scriptsize ARNESS}-Z{\scriptsize ERO}}}{\text{H{\scriptscriptstyle ARNESS}-Z{\scriptscriptstyle ERO}}}{\text{HARNESS-ZERO}}}HARNESS-ZERO spans specialized tool use expressed in the student's native action space, behavioral patterns enforced by middleware code, and accumulated knowledge and reusable experience supplied by skills and memory. The source and target harnesses can differ in both action space and available information, making trajectories collected under the source harness unsuitable for direct imitation under the target harness.
Our solution to this challenge is agent-as-harness, which wraps the student agent with a harnessing agent at its response boundary. We denote the student's fixed operating harness as the target harness hhh, the evolved student-side harness (via, e.g., meta-harness ([11])) as h⋆h^\starh⋆, and its adaptation for the harnessing agent as the private reference harness K\mathcal KK. As one example of this adaptation, a middleware rule in h⋆h^\starh⋆ that blocks risky actions becomes a review-time warning in K\mathcal KK, activated when the student proposes such an action.
During data collection, the harnessing agent uses K\mathcal KK to review each proposed response before it is executed or added to the trajectory. It passes sound proposals unchanged and otherwise makes the smallest coherent correction, expressed as a complete response in the student's action space. The accepted response is executed through hhh, and the resulting observation is appended to the student's trajectory. The harnessing agent cannot inspect the student sandbox or consult hidden task solutions, and its private review discussion remains outside the student-visible trajectory. HARNESS-ZERO{\mathchoice{\text{H{\scriptsize ARNESS}-Z{\scriptsize ERO}}}{\text{H{\scriptsize ARNESS}-Z{\scriptsize ERO}}}{\text{H{\scriptscriptstyle ARNESS}-Z{\scriptscriptstyle ERO}}}{\text{HARNESS-ZERO}}}HARNESS-ZERO applies supervised fine-tuning (SFT) to these reviewed rollouts to internalize the demonstrated behaviors in model parameters.
We evaluate HARNESS-ZERO{\mathchoice{\text{H{\scriptsize ARNESS}-Z{\scriptsize ERO}}}{\text{H{\scriptsize ARNESS}-Z{\scriptsize ERO}}}{\text{H{\scriptscriptstyle ARNESS}-Z{\scriptscriptstyle ERO}}}{\text{HARNESS-ZERO}}}HARNESS-ZERO across three domains: spreadsheet-based knowledge work (SpreadsheetBench Verified), multi-application tool use (AppWorld), and scientific reasoning (USPTO Retrosynthesis). We instantiate hhh as a fixed, minimal mini-SWE-agent ([21]) with one Bash
execute tool. We first evaluate agent-as-harness on frontier models without training and find that it outperforms code-as-harness (81.1% vs. 78.1%, averaged across six benchmark–model settings). We then evaluate harness distillation on Qwen3.5-9B and remove h⋆h^\starh⋆, K\mathcal KK, and the harnessing agent at deployment. Under hhh alone, HARNESS-ZERO{\mathchoice{\text{H{\scriptsize ARNESS}-Z{\scriptsize ERO}}}{\text{H{\scriptsize ARNESS}-Z{\scriptsize ERO}}}{\text{H{\scriptscriptstyle ARNESS}-Z{\scriptscriptstyle ERO}}}{\text{HARNESS-ZERO}}}HARNESS-ZERO raises macro-average performance from 23.3% to 44.3%, exceeding the 41.7% obtained by the base model with h⋆h^\starh⋆ still attached. Controlled ablations further show that harness-guided review produces substantially better distilled performance than alternative supervision sources, including direct trajectories from a stronger model, trajectories generated under h⋆h^\starh⋆, review without K\mathcal KK, and review given only the task answer (30% vs. 3–15%). Behavioral analysis also finds that the distilled model recovers most behaviors induced by h⋆h^\starh⋆ but absent from the base model (82.3% recovery on average across 28 patterns).Contributions. ❶ We formulate agent harness distillation and present HARNESS-ZERO{\mathchoice{\text{H{\scriptsize ARNESS}-Z{\scriptsize ERO}}}{\text{H{\scriptsize ARNESS}-Z{\scriptsize ERO}}}{\text{H{\scriptscriptstyle ARNESS}-Z{\scriptscriptstyle ERO}}}{\text{HARNESS-ZERO}}}HARNESS-ZERO, which transfers behaviors induced by an optimized harness into model parameters for deployment under a fixed target harness. ❷ We introduce agent-as-harness, which translates guidance from an optimized harness into executable supervision at the student's response boundary, enabling imitation learning across different harness action spaces. ❸ Experimental results show that agent-as-harness can outperform code-as-harness, that HARNESS-ZERO{\mathchoice{\text{H{\scriptsize ARNESS}-Z{\scriptsize ERO}}}{\text{H{\scriptsize ARNESS}-Z{\scriptsize ERO}}}{\text{H{\scriptscriptstyle ARNESS}-Z{\scriptscriptstyle ERO}}}{\text{HARNESS-ZERO}}}HARNESS-ZERO retains the gains of optimized harnesses after they are removed, and that the distilled model recovers harness-induced behaviors.
2. Related work
Code-as-harness and its optimization. Harnesses mediate the interaction between LLMs and their environments through tools, context management, control flow, and persistent state ([1, 2, 22]). Modern coding agents such as Claude Code ([23]), Codex ([24]), Kimi Code ([25]), Pi ([26]), and OpenCode ([27]) embody different design philosophies within this space. Natural-Language Agent Harnesses represent run-level policies as editable documents, which an agent interprets into actions ([28]). Automatic, training-free optimization has expanded from prompts ([29, 30, 31]) and context ([17, 5]) to workflows ([32, 33]) and harnesses ([11, 12, 13]). [34] train a model to generate per-task harnesses just in time. Harness design can also support test-time strong-to-weak transfer, in which a stronger model builds an inference-time harness for a fixed weaker model ([35]). Complementing this line of work, HARNESS-ZERO{\mathchoice{\text{H{\scriptsize ARNESS}-Z{\scriptsize ERO}}}{\text{H{\scriptsize ARNESS}-Z{\scriptsize ERO}}}{\text{H{\scriptscriptstyle ARNESS}-Z{\scriptscriptstyle ERO}}}{\text{HARNESS-ZERO}}}HARNESS-ZERO introduces agent-as-harness, which can outperform code-as-harness for frontier models and supports distilling optimized harness behavior into model weights.
Co-evolution of model and harness. Several approaches combine harness optimization with parameter updates. One line alternates between the two, using revised harnesses to generate data for the next model update and updated models to motivate further harness search ([36, 37, 38, 7]). Other methods optimize the two jointly, treating model–harness compatibility as the objective and adapting the agent harness together with the policy trained from its trajectories ([39, 40, 41, 42]). These studies show that harness and weight updates can reinforce each other, but the resulting gains may remain coupled to the harness. A controlled coding-agent study supports this concern ([43]): changing the evaluation harness affected performance more than the training method, and training with feedback collected across harnesses did not improve transfer to a held-out minimal ReAct harness. HARNESS-ZERO{\mathchoice{\text{H{\scriptsize ARNESS}-Z{\scriptsize ERO}}}{\text{H{\scriptsize ARNESS}-Z{\scriptsize ERO}}}{\text{H{\scriptscriptstyle ARNESS}-Z{\scriptscriptstyle ERO}}}{\text{HARNESS-ZERO}}}HARNESS-ZERO instead uses optimized harnesses to guide a temporary harnessing agent and distills the resulting behavior into model weights, internalizing harness gains without retaining or routing the harness collection at deployment.
Distillation from privileged guidance. EvoHarness-RL provides early evidence of harness internalization on ALFWorld, where its trained agent learns to manage external state and makes fewer, more selective harness calls ([44]). OPHSD distills privileged outputs produced by sequential draft–verify and plan–solve LLM workflows into a standalone model ([45]). Together, these studies provide preliminary evidence that models can absorb harness-induced behavior into their parameters. However, each addresses only an individual mechanism rather than a complete tool-using agent harness; EvoHarness-RL retains the external workspace at deployment, and OPHSD does not involve an interactive agent loop. HARNESS-ZERO{\mathchoice{\text{H{\scriptsize ARNESS}-Z{\scriptsize ERO}}}{\text{H{\scriptsize ARNESS}-Z{\scriptsize ERO}}}{\text{H{\scriptscriptstyle ARNESS}-Z{\scriptscriptstyle ERO}}}{\text{HARNESS-ZERO}}}HARNESS-ZERO introduces a general method for distilling complete tool-using agent harnesses into a model that runs under a minimal target harness at deployment.
3. Harness-Zero
HARNESS-ZERO{\mathchoice{\text{H{\scriptsize ARNESS}-Z{\scriptsize ERO}}}{\text{H{\scriptsize ARNESS}-Z{\scriptsize ERO}}}{\text{H{\scriptscriptstyle ARNESS}-Z{\scriptscriptstyle ERO}}}{\text{HARNESS-ZERO}}}HARNESS-ZERO trains a student model to reproduce behaviors induced by an evolved harness. The method has three stages (Figure 1). First (Section 3.1), we evolve a student-side harness on training tasks and adapt it into a private reference harness for a separate harnessing agent. Second (Section 3.2), the harnessing agent wraps the student's response loop for training trajectory collection. Third (Section 3.3), we apply SFT to reviewed trajectories jointly produced by the student and the harnessing agent. At deployment, the distilled student aims to retain the evolved harness's gains under the target harness alone.
3.1 Harness evolution and adaptation
The first stage builds the domain-specific guidance used during training trajectory collection. We denote the fixed target harness by hhh, the evolved student-side harness by h⋆h^\starh⋆, and the private reference harness adapted from it by K\mathcal KK. Given training tasks Dtrain\mathcal D_{\mathrm{train}}Dtrain, we write the two steps as
Harness evolution. We evolve h⋆h^\starh⋆ on training tasks. This step can use existing automated harness-optimization methods ([11, 12, 13]). In our implementation, a simple skill-guided evolution loop analyzes recurring failures and updates the domain harness. The resulting h⋆h^\starh⋆ follows the DeepAgents abstraction ([46]), which includes tools, middleware, skills, and memory. Appendix A describes the three-round procedure and presents the evolution skill.
Harness adaptation. The evolved student-side harness h⋆h^\starh⋆ is designed to act directly around the student. Tools extend its action space, middleware modifies or blocks its execution loop, and skills and memory instruct the student. We adapt these components into the reference harness K\mathcal KK used by the harnessing agent; the process can be automated by an agent. The adaptation is relative to hhh, because any correction the harnessing agent constructs from K\mathcal KK must ultimately be executed by the student in the target harness's action space. In general, the adaptation preserves the harness components' intended behavior while changing their audience and enforcement point. Tools become specifications for constructing student-native equivalents; student-side middleware becomes review middleware that privately alerts the harnessing agent when the corresponding condition is met; and skills and memory become diagnostic criteria and intervention guidance. The adaptation may also produce a domain prompt appended to the harnessing agent's system prompt, stating the domain's review policy. Appendix B gives detailed examples of such adaptation.
3.2 Agent-as-harness
Let πθh\pi_\theta^hπθh denote the response policy induced when a model with parameters θ\thetaθ operates under hhh, and let Th(c,y)T_h(c,y)Th(c,y) denote the transition function that executes response yyy through hhh from context ccc and returns the next student-visible context.
Code-as-harness, after harness evolution, runs the base student with parameters θ0\theta_0θ0 directly under h⋆h^\starh⋆ ([2]):
Agent-as-harness instead runs the student under hhh, with a harnessing agent wrapping it at its response boundary. The harnessing agent intercepts each proposed response before execution and either passes it or replaces it. It uses K\mathcal KK for private guidance on when to intervene and how to construct a replacement; K\mathcal KK neither acts on the environment nor enters the student-visible context. The target harness hhh defines the student's action space and executes the accepted response. The student and harnessing agent may use the same underlying model, with different instructions and context for their respective roles.
At turn ttt, let ctc_tct denote the student's visible context: the task, previous accepted responses, and resulting environment observations. The student policy πS\pi_SπS proposes a response yty_tyt, which the harnessing policy πH\pi_HπH reviews using ctc_tct, the guidance in K\mathcal KK, and its private history r<tr_{<t}r<t:
where dtd_tdt is the review decision, ztz_tzt is a complete replacement response valid under hhh, and r<tr_{<t}r<t contains earlier review exchanges and the harnessing agent's prior file-system reads from K\mathcal KK. A single harnessing-agent session spans the entire student rollout. At each review, it receives the student-visible events added since the previous review and the current unexecuted proposal. Only the accepted response y~t\tilde{y}_ty~t enters the student-visible trajectory. Any actions it contains are then executed through hhh, and their observations become part of ct+1c_{t+1}ct+1. The rejected proposal and private review remain outside this trajectory. This process realizes source-harness guidance as a target-harness trajectory incrementally, with each correction conditioned on the student's current interaction history. Appendix C specifies the review process and gives the harnessing agent's system prompt.
Intervention policy and constraints. Guidance from K\mathcal KK steers rollouts toward behaviors and states that an unaided student under hhh may not reach. To facilitate SFT, the harnessing agent minimizes changes to the student's proposals. It passes sound proposals unchanged. When intervention is necessary, it makes the smallest coherent correction needed to follow K\mathcal KK's guidance and preserves the rest of the proposal whenever possible. Each replacement is a complete response valid under hhh that continues from the current student-visible state.
The harnessing agent's privileged access is limited to reading K\mathcal KK. It cannot inspect hidden solutions or verifier feedback, nor can it access environment state outside ctc_tct. Any additional task evidence must therefore be obtained by proposing an action available under hhh. For example, it can replace a premature completion with code that checks the student's work. Executing the code through hhh adds both the check and its result to the student-visible trajectory, grounding the intervention in student-observable evidence. Together, these intervention and grounding constraints make the reviewed trajectories directly usable to fine-tune the student for operation under hhh.
3.3 Learning from reviewed trajectories
We perform imitation learning under the target harness by applying SFT to the accepted responses in reviewed rollouts Equation (3), including both unchanged student proposals and harness-guided replacements. Replacements should be self-contained responses written from the student's perspective. Because they are generated within the private review context, they may inadvertently include reviewer-perspective reasoning about the student's proposal or the review decision. We mask such reasoning from the loss; Appendix D details data collection and the filtering rule. Let Dreview\mathcal D_{\mathrm{review}}Dreview denote the retained trajectories, we optimize:
Here, TτT_\tauTτ is the number of accepted response turns in trajectory τ\tauτ. At deployment, the distilled policy acts directly under the same target harness:
The deployed system is therefore (πθ^,h)(\pi_{\hat{\theta}},h)(πθ^,h), without h⋆h^\starh⋆, K\mathcal KK, or the harnessing agent. The training objective is for πθ^\pi_{\hat{\theta}}πθ^ to reproduce under hhh the behavior patterns induced by h⋆h^\starh⋆.
4. Experiments
We first compare agent-as-harness with code-as-harness at inference time, then test whether HARNESS-ZERO{\mathchoice{\text{H{\scriptsize ARNESS}-Z{\scriptsize ERO}}}{\text{H{\scriptsize ARNESS}-Z{\scriptsize ERO}}}{\text{H{\scriptscriptstyle ARNESS}-Z{\scriptscriptstyle ERO}}}{\text{HARNESS-ZERO}}}HARNESS-ZERO can distill an optimized harness into model weights.
4.1 Experimental setup
Tasks and splits. We evaluate three task domains. (1) SpreadsheetBench Verified contains 400 real-world spreadsheet-manipulation tasks ([47]); we use 300 for harness evolution and training data collection and hold out 100 for evaluation. (2) AppWorld evaluates interactive tool use across simulated applications ([48]); we merge its official train and development sets into a 147-task training split and hold out the 168
test_normal tasks, grouped into 56 three-task scenarios. (3) USPTO Retrosynthesis covers single-step precursor prediction from the USPTO reaction corpus ([49, 50]); we use a 500-task training split balanced across its ten reaction classes and a disjoint 100-task test split. All three run in the Harbor framework ([51, 52]). We report single-run task success (pass@1, %) for SpreadsheetBench and USPTO and scenario goal completion (SGC, %) for AppWorld.Models and training. The target harness hhh is a minimal mini-SWE-agent-style harness ([21]) with a fixed system prompt and a single Bash execution tool. The training-free experiments use GPT-5.6 Sol ([53]) and DeepSeek-V4-Pro ([54]). The distillation experiments use Qwen3.5-9B ([55]) as the base model and GPT-5.6 Sol as the harnessing agent; harness evolution uses Kimi K3 ([56]) under Kimi Code. Reasoning is enabled for all models, with reasoning effort set to high when applicable. After rollout collection and filtering, the training data comprise 487 rollouts for SpreadsheetBench, 282 for AppWorld, and 500 for USPTO. We perform LoRA SFT on Qwen3.5-9B for two epochs using the Tinker recipe ([57]). Appendix E gives the student prompt, the model access routes, and the full training configuration.
4.2 Evaluating agent-as-harness
We first compare agent-as-harness against code-as-harness at inference time, with no parameter updates. Table 1 varies two factors: whether the evolved harness is available, and whether it reaches the student as code wrapped around it or as a harnessing agent reviewing its responses. The two code-as-harness conditions use no harnessing agent: mini-SWE-agent runs the student under hhh alone, and meta-harness mounts h⋆h^\starh⋆ on top of it. The two agent-as-harness conditions keep the student under hhh and add a harnessing agent that consults either an empty reference harness (w/o evolved) or K\mathcal KK adapted from h⋆h^\starh⋆ (w/ evolved). In both, the same model plays student and harnessing agent, so the gains cannot come from a stronger supervising model. On the frontier models we reuse the h⋆h^\starh⋆ evolved on Qwen3.5-9B.
Obs.❶ With evolved harness, agent-as-harness outperforms code-as-harness on average. Across the six settings in Table 1, agent-as-harness with the adapted K\mathcal KK averages 81.1%, against 78.1% for meta-harness and 68.6% for mini-SWE-agent. With an empty K\mathcal KK it averages only 69.2%, so review alone explains little of the gain. Beyond these benchmark numbers, agent-as-harness also offers better adaptability across model updates. A code harness encodes assumptions about how a model should act, and as capabilities change those assumptions go stale, forcing the harness to be re-adapted for each new model ([35, 15]). Agent-as-harness moves that adaptation into inference: the harnessing agent interprets K\mathcal KK against the current trajectory and decides when and how to intervene, so the guidance stays reusable and a stronger harnessing model directly improves how it is applied. While this approach requires a sufficiently capable harnessing agent (Appendix F), we expect its advantage over fixed code harnesses to widen as foundation models continue to improve.
4.3 Evaluating agent harness distillation
We next test whether the behavior induced by the optimized harness h⋆h^\starh⋆ can be retained after distillation, when h⋆h^\starh⋆, reference harness K\mathcal KK, and harnessing agent are removed. Table 2 compares the base model under hhh, the base model with h⋆h^\starh⋆ mounted, and the distilled model under hhh. As reference points, we also run the base model under two general-purpose harnesses: DeepAgents ([46]), the abstraction on which h⋆h^\starh⋆ is built (Section 3.1), and Claude Code ([23]).
Obs.❷ Distillation raises the base model's macro-average performance by 21.0 points and surpasses h⋆h^\starh⋆. HARNESS-ZERO{\mathchoice{\text{H{\scriptsize ARNESS}-Z{\scriptsize ERO}}}{\text{H{\scriptsize ARNESS}-Z{\scriptsize ERO}}}{\text{H{\scriptscriptstyle ARNESS}-Z{\scriptscriptstyle ERO}}}{\text{HARNESS-ZERO}}}HARNESS-ZERO raises the macro average from 23.3% to 44.3%, a 21.0-point absolute gain and a 90.1% relative improvement. It also exceeds the 41.7% macro average of the base model equipped with h⋆h^\starh⋆. Neither general-purpose harness benefits the base model. On one hand, a 9B model handles their larger, generic tool suites and extended context poorly. On the other hand, domains such as USPTO and AppWorld demand domain-specific tooling and constraints present in h⋆h^\starh⋆ (such as molecular validation tools for USPTO), rendering generic tools beyond basic bash largely unhelpful and distracting.
Obs.❸ Procedural harness behavior is easier to internalize than deep domain knowledge. On SpreadsheetBench and AppWorld, HARNESS-ZERO{\mathchoice{\text{H{\scriptsize ARNESS}-Z{\scriptsize ERO}}}{\text{H{\scriptsize ARNESS}-Z{\scriptsize ERO}}}{\text{H{\scriptscriptstyle ARNESS}-Z{\scriptscriptstyle ERO}}}{\text{HARNESS-ZERO}}}HARNESS-ZERO under hhh alone surpasses h⋆h^\starh⋆. Their harnesses mainly encode recurring procedures for state inspection, targeted changes, and verification, which reviewed trajectories can demonstrate directly. On USPTO, HARNESS-ZERO{\mathchoice{\text{H{\scriptsize ARNESS}-Z{\scriptsize ERO}}}{\text{H{\scriptsize ARNESS}-Z{\scriptsize ERO}}}{\text{H{\scriptscriptstyle ARNESS}-Z{\scriptscriptstyle ERO}}}{\text{HARNESS-ZERO}}}HARNESS-ZERO improves over the base model (30.0% vs. 12.0%) but trails h⋆h^\starh⋆ (38.0%). There, h⋆h^\starh⋆ also supplies reaction priors, candidate-generation logic, and executable SMILES validation; transferring this knowledge and functionality may require broader pretraining or mid-training coverage, or more distillation trajectories.
4.4 Comparing alternative distillation signals
To ablate the agent-as-harness recipe and understand what makes it effective, we compare against alternative training trajectory sources on USPTO. We vary the rollout generator, the use of h⋆h^\starh⋆, and the private context available to the harnessing agent. The three direct-distillation baselines collect rollouts from the teacher (GPT-5.6 Sol, used in the harnessing agent of HARNESS-ZERO{\mathchoice{\text{H{\scriptsize ARNESS}-Z{\scriptsize ERO}}}{\text{H{\scriptsize ARNESS}-Z{\scriptsize ERO}}}{\text{H{\scriptscriptstyle ARNESS}-Z{\scriptscriptstyle ERO}}}{\text{HARNESS-ZERO}}}HARNESS-ZERO) under hhh, the teacher with h⋆h^\starh⋆ mounted, and the base student with h⋆h^\starh⋆ mounted. The two harnessing-agent controls keep the base student under hhh and give the harnessing agent either an empty K\mathcal KK or the oracle answer.
All conditions start from Qwen3.5-9B, use the same 500-task collection split and training recipe, and evaluate the 2-epoch checkpoint under hhh on the test set. Trajectories are retained after the same structural and privacy checks. For each condition, Table 3 reports the source's success rate on the 500 collection tasks and the test pass@1 of the resulting student under hhh.
Obs.❹ Effective supervision comes from harness-guided review, not from stronger demonstrations alone. Directly fine-tuning on GPT-5.6 Sol trajectories leaves the student at the 12% base result, even though that source succeeds on 52.0% of the collection tasks. Likewise, trajectories collected by the base student under h⋆h^\starh⋆ yield 12% after SFT. This comparison isolates the importance of starting from the student's trajectory and translating harness guidance into corrections compatible with its current state and target action space.
Obs.❺ Procedural guidance provides better distillation supervision than answer access or generic review. Review with an empty K\mathcal KK reaches only 11%. Providing oracle answers raises collection success to 98.6%, yet the distilled model reaches only 15%. By comparison, K\mathcal KK contains no task answers and reaches a lower collection success of 59.4%, yet the student distilled from its trajectories reaches 30% pass@1. Collection success therefore does not predict distillation value. Oracle access encourages non-generalizable shortcut corrections. By contrast, K\mathcal KK instills reusable procedural behavior (e.g., reasoning from reaction templates, proposing candidates, and systematically validating reactant sets) that the student can execute independently under hhh.
Obs.❻ Executing h⋆h^\starh⋆ during collection does not make its behavior transferable. Mounting h⋆h^\starh⋆ improves GPT-5.6 Sol's collection success from 52.0% to 62.0%, but the resulting distilled student falls from 12% to 3%. The student-generated counterpart also succeeds on 39.4% of collection tasks but returns to 12% after SFT. The failure is consistent with action-space mismatch. The model distilled from teacher trajectories under h⋆h^\starh⋆ extensively attempts unavailable harness-tool calls, and many trials thus exhaust the turn limit. HARNESS-ZERO{\mathchoice{\text{H{\scriptsize ARNESS}-Z{\scriptsize ERO}}}{\text{H{\scriptsize ARNESS}-Z{\scriptsize ERO}}}{\text{H{\scriptscriptstyle ARNESS}-Z{\scriptscriptstyle ERO}}}{\text{HARNESS-ZERO}}}HARNESS-ZERO avoids this mismatch by expressing each correction through hhh.
4.5 Behavioral internalization
To evaluate harness distillation at a finer granularity, we measure how well the distilled model internalizes the behavioral patterns encoded in h⋆h^\starh⋆. We first translate patterns of enabled tools, middleware, memory, and skills in h⋆h^\starh⋆ into trajectory detectors. For each detector, we compare the base model's trajectories under h⋆h^\starh⋆ and hhh on the same test task. We select tasks where the pattern appears only under h⋆h^\starh⋆ and keep a pattern only when at least 10 tasks meet this criterion. This yields 18 harness-exclusive patterns on SpreadsheetBench, 6 on USPTO, and 4 on AppWorld; Appendix G details the mining procedure and explains every pattern. For each pattern, its recovery rate is the fraction of the selected tasks on which HARNESS-ZERO{\mathchoice{\text{H{\scriptsize ARNESS}-Z{\scriptsize ERO}}}{\text{H{\scriptsize ARNESS}-Z{\scriptsize ERO}}}{\text{H{\scriptscriptstyle ARNESS}-Z{\scriptscriptstyle ERO}}}{\text{HARNESS-ZERO}}}HARNESS-ZERO exhibits the same behavior.
Obs.❼ HARNESS-ZERO{\mathchoice{\text{H{\scriptsize ARNESS}-Z{\scriptsize ERO}}}{\text{H{\scriptsize ARNESS}-Z{\scriptsize ERO}}}{\text{H{\scriptscriptstyle ARNESS}-Z{\scriptscriptstyle ERO}}}{\text{HARNESS-ZERO}}}HARNESS-ZERO internalizes behaviors contributed by h⋆h^\starh⋆. Averaged over the 28 patterns in Table 4, HARNESS-ZERO{\mathchoice{\text{H{\scriptsize ARNESS}-Z{\scriptsize ERO}}}{\text{H{\scriptsize ARNESS}-Z{\scriptsize ERO}}}{\text{H{\scriptscriptstyle ARNESS}-Z{\scriptscriptstyle ERO}}}{\text{HARNESS-ZERO}}}HARNESS-ZERO recovers 82.3% of the harness-exclusive behaviors. Recovery spans memory, skill, tool, and middleware sources, covering both model-visible guidance and executable components. Each pattern is scored only on the tasks selected for it, where the base model exhibits the behavior under h⋆h^\starh⋆ but never under hhh, so the base model scores 0% on every pattern by construction. These recovery rates therefore measure behavior that distillation adds, providing direct evidence that HARNESS-ZERO{\mathchoice{\text{H{\scriptsize ARNESS}-Z{\scriptsize ERO}}}{\text{H{\scriptsize ARNESS}-Z{\scriptsize ERO}}}{\text{H{\scriptscriptstyle ARNESS}-Z{\scriptscriptstyle ERO}}}{\text{HARNESS-ZERO}}}HARNESS-ZERO internalizes behavior induced by h⋆h^\starh⋆.
5. Conclusion and discussion
HARNESS-ZERO{\mathchoice{\text{H{\scriptsize ARNESS}-Z{\scriptsize ERO}}}{\text{H{\scriptsize ARNESS}-Z{\scriptsize ERO}}}{\text{H{\scriptscriptstyle ARNESS}-Z{\scriptscriptstyle ERO}}}{\text{HARNESS-ZERO}}}HARNESS-ZERO turns an evolved student-side harness h⋆h^\starh⋆ into training supervision for a model operating under a fixed target harness hhh. Guided by the adapted reference harness K\mathcal KK, a harnessing agent rewrites the student's proposals before execution, producing training trajectories compatible with hhh. Across three domains, agent-as-harness outperforms code-as-harness on frontier models (81.1% vs. 78.1% on average), and ablations identify harness-guided review as the most effective of the tested supervision sources. After SFT, the student under hhh alone raises macro-average task success from 23.3% to 44.3%, surpassing the 41.7% achieved with h⋆h^\starh⋆, and recovers h⋆h^\starh⋆-specific behaviors at an average rate of 82.3% over 28 patterns.
Limitations. The method depends on a capable harnessing model. With weaker models, review can become harmful and agent-as-harness loses its advantage over code-as-harness (Appendix F). Reviewing every proposal also increases collection cost; each step requires an additional model call that processes the trajectory and K\mathcal KK, raising mean USPTO latency to 2.4×2.4\times2.4×. This overhead applies during trajectory collection and is absent after distillation. SFT may also fail to fully internalize deep domain knowledge encoded by h⋆h^\starh⋆. In addition, some harness mechanisms, e.g., context management, are not fully expressible as student responses out of the box. HARNESS-ZERO{\mathchoice{\text{H{\scriptsize ARNESS}-Z{\scriptsize ERO}}}{\text{H{\scriptsize ARNESS}-Z{\scriptsize ERO}}}{\text{H{\scriptscriptstyle ARNESS}-Z{\scriptscriptstyle ERO}}}{\text{HARNESS-ZERO}}}HARNESS-ZERO therefore does not remove the need for a harness, but narrows what that harness must provide, and we expect a minimal one to suffice as harness distillation improves.
Future work. A harnessing agent often faces a counterfactual prediction problem. It must anticipate how executing the student's proposal would change the environment and whether an intervention would produce a better trajectory. Effective harnessing therefore depends on an accurate model of agent–environment dynamics. A promising model-level direction is to train stronger agent world models ([58]) and specialize their predictive capabilities for harnessing decisions. At the framework level, the current design reviews every student proposal and restricts intervention to passing or replacing the full response. This incurs unnecessary calls on sound proposals and provides only coarse-grained control. Future work could use proxy signals to invoke review selectively and support finer-grained mechanisms, such as token insertion and latent-space steering. The training method can also be improved. For example, each replacement pairs a rejected and a preferred response at the same state, so preference learning could use comparison signals that response-level SFT discards.
Overall, we hope HARNESS-ZERO{\mathchoice{\text{H{\scriptsize ARNESS}-Z{\scriptsize ERO}}}{\text{H{\scriptsize ARNESS}-Z{\scriptsize ERO}}}{\text{H{\scriptscriptstyle ARNESS}-Z{\scriptscriptstyle ERO}}}{\text{HARNESS-ZERO}}}HARNESS-ZERO helps pave the way for a new paradigm of agent harness and recursive self-improvement. Distilling many domain- and task-specific harnesses into a shared model would let behaviors and knowledge developed across agent systems accumulate in model parameters instead of remaining fragmented across external scaffolds. Harness development could then become a scalable source of training signal, with better models building better harnesses and each harness returning its gains to the weights.
AI Use Statement
We used AI assistants in two roles. First, to check grammar and to polish text the authors had written. Second, for routine coding assistance during implementation. The method, the experimental design, and every implementation decision affecting the reported results were made by the authors, who verified all AI-assisted output and take full responsibility for this paper.
Appendix
A.1 Harness-Zero: Harness Distillation via Agent-as-Harness (Appendix)
A. Skill-guided harness evolution
A.1 Evolution methods
We evolve one shared student-side harness for each benchmark. The process, also known as meta-harness, begins with rollouts under hhh on the training split. In each round, a swarm of analysis agents examines the tasks that failed in the preceding round and proposes reusable changes to tools, middleware, skills, or memory. A main evolution agent consolidates these proposals into the shared harness and evaluates the same model with that harness attached. Tasks that pass leave the evolution process, while the remaining failures form the next round. We run this process for three rounds.
The evolution objective is performance under the student-side harness itself. Evaluation therefore uses the target student agent with the evolved harness attached and no harnessing-agent intervention. Proposed components must address failure patterns shared across tasks. Algorithm 1 shows the skill followed by the evolution agent.
---
name: student-harness-evolve
description: Generic workflow for evolving a shared student harness -- for a benchmark's Harbor-format train task set, take bare miniswe rollouts as the baseline, handle only the tasks the previous round got wrong, use a swarm (default 10 coder subagents) to analyze rollout trajectories in parallel, have the main agent write the shared student harness (tools / middleware / skills / memory) under 'harness_bank/<domain>/', and iterate for 3 rounds with "miniswe eval rollouts loading that harness" as the evaluator. The evolved artifact can later be rewritten into a teacher-side harness.
---
# Shared Student Harness Evolution Guide
## What this workflow does
For a benchmark's train task set, evolve **one** deliverable: a **student-side harness shared by all tasks**, kept in a per-benchmark bank ('harness_bank/<domain>/'):
- 'tools/' -- prebuilt tools (real StructuredTools loaded into miniswe alongside 'execute');
- 'middlewares/' -- student-side langchain AgentMiddleware;
- 'skills/<name>/SKILL.md' + 'memory.md' -- skill and memory material, uploaded into the sandbox and exposed to the student via deepagents' SkillsMiddleware / MemoryMiddleware.
**The evaluator** is a bare miniswe rollout with the harness loaded ('ahd.harness:AHDMinisweAgent', teacher passthrough, i.e. no teacher intervention; the bank is mounted via 'student_harness_dir'). It measures "does this student harness help the model itself". A campaign runs **3 rounds** on the rhythm: "select the tasks the previous round got wrong -> parallel swarm analysis -> main agent writes the bank -> eval rollout on exactly those failed tasks -> attribution". Tasks already solved count as "the harness is good enough for them" and are not rolled out again in later rounds.
## Campaign parameters (set at each instantiation)
| Parameter | Meaning |
|---|---|
| '<DOMAIN>' | Domain name (determines the bank path 'harness_bank/<domain>/'; for the student side, prefer a '<benchmark>_student' suffix to distinguish it from teacher-side banks) |
| '<TASKS>' | Harbor task path + task-list file (e.g. '-p data/spreadsheetbench-verified', task list 'experiments/spreadsheetbench_train300.txt') |
| '<MODEL>' | Evaluation model (e.g. 'openrouter:qwen/qwen3.5-9b', reasoning explicitly enabled) |
| '<BASELINE_JOBS>' | Job directory of the baseline (bare, no harness) rollout; the input for round-1 failure analysis. If it does not exist, run the baseline first |
| '<N_AGENTS>' | Swarm concurrency, at most 10 subagents at a time; all of the round's failed tasks are distributed among these subagents, with no per-subagent task cap |
## Bank structure and integration
```
harness_bank/<domain>/
|-- registry.py # single entry point; API contract below
|-- manifest.json # enablement manifest: tools/middlewares/skills lists + memory switch; toggle components only here
|-- memory.md # accumulated checklist clauses (English, no source attribution, grouped by topic)
|-- tools/ # *.py, each exposing make_tool(backend) -> StructuredTool
|-- middlewares/ # *.py, each exposing make_middleware() -> AgentMiddleware
|-- skills/ # <name>/SKILL.md, deepagents skill format (progressive disclosure)
`-- README.md # bank description and provenance
```
- **Registry API contract** (consumed by the loader under exactly these names): 'TOOL_FACTORIES' / 'MIDDLEWARE_FACTORIES' (dict, name -> '"module:function"' import string, which the loader resolves into a factory). 'catalog_text()' (component listing rendering) is only for logs and audits; the loader does not consume it. Component docstrings are the catalog source.
- **Loading semantics** (the contract of 'student_harness_dir', implemented in 'src/ahd/student_harness.py' / 'src/ahd/student.py'):
1. tools are appended to miniswe's tools list (after 'execute');
2. the 'skills/' directory and memory.md are uploaded into the sandbox ('/opt/ahd/harness/skills/', '/opt/ahd/harness/memory.md'): skills are progressively disclosed through deepagents' 'SkillsMiddleware' (only the index goes into the system prompt; the student reads full text via 'execute cat'), and memory.md is injected into the system prompt through 'MemoryMiddleware'; SkillsMiddleware and MemoryMiddleware run **before** the review middleware so the teacher sees the same context as the student;
3. the bank's middleware is appended after the review middleware, followed by 'StudentRequestCaptureMiddleware' (records the final model request, for audit);
4. **bank bytes are hashed into 'bank_sha256'** -- the rollout log writes 'student_harness_bank.txt' recording the mounted bank's hash, for provenance only, with no assertions;
5. when 'student_harness_dir' is not passed, behavior is exactly the status quo (bare compatible).
- **The student core is not evolvable**: the 'execute' tool, the student.md base prompt, the review middleware, the turn limit, and the tool_call adapter are fixed layers; evolution happens only in bank content.
## Per-round workflow
### Phase 0: determine this round's failed-task set
- Round 1: select tasks whose reward did not pass from the '<BASELINE_JOBS>' results (if the baseline does not exist, first run bare rollouts over the train task list with '<MODEL>').
- Rounds 2/3: select tasks whose reward did not pass from the previous round's harnessed eval rollout.
- This round's swarm, bank edits, and eval rollout cover only these failed tasks. Tasks already solved leave the campaign.
- If the failed-task set is empty, end the campaign early; completed rounds still count as valid results.
### Phase 1: swarm (at most 10 coder subagents at a time; analyze and propose changes only)
Launch with AgentSwarm, item = this round's failed-task group. Each subagent does two things for its assigned tasks:
1. **Failure analysis -> 'experiments/<domain>_evolution/round<N>/notes/<TASK>.md'**
- Round 1: read the task's bare trajectory in '<BASELINE_JOBS>' ('<job>/<TASK>__*/agent/trajectory.json', 'llm_calls.jsonl', 'verifier/'), and understand how the agent worked and the failure class (insufficient exploration / mechanism misused / miscalculation / deliverable missing or misplaced / missing verification discipline).
- Rounds 2/3: read the previous round's harnessed trajectory and attribute per the Attribution checklist section.
2. **Propose bank changes**: per Evolution discipline, decide whether the failure should be addressed by generalizing an existing component, adding a new component, or editing prompt/memory material; write the proposal and its rationale in the notes (it must argue "this helps a class of tasks"); do not write to 'harness_bank/' directly.
Subagent forbidden zone: do not modify 'harness_bank/', 'src/', or 'data/'; do not run rollouts; do not perform git operations.
**WARNING -- component hard rule:** middleware hooks run on the harbor event loop; sandbox probing may only use 'await backend.aexecute(...)'; never call the synchronous 'backend.execute(...)' inside an async hook -- internally it is 'run_coroutine_threadsafe(...).result()', which deadlocks the entire event loop when called on the loop thread (0% CPU, all trials frozen, even trial-level timeouts cannot fire). Tool function bodies run on worker threads and are not subject to this restriction; message-level checks in middleware are always the safest choice.
### Phase 2: main-agent write-up and acceptance
1. **Consolidated writing**: read all of the round's notes, resolve duplicate or conflicting proposals, and have the main agent uniformly modify the bank (tools / middleware / skills / memory.md / manifest.json).
2. **Reconciliation**: grep every reference name in manifest.json against the registered names in the registry; stale references must be zeroed out (otherwise loading fails outright).
3. **Smoke test**: import smoke ('from harness_bank.<domain>.registry import ...' + render the catalog); 'load_student_harness(bank_dir)' resolves successfully (all manifest references hit, 'skills/<name>/SKILL.md' files all present); 'pytest tests/ -q' shows no regressions.
4. **Eval rollout**: submit a rollout of bare miniswe with the bank loaded, using this round's failed-task IDs as the exact include set (teacher passthrough, same model as the baseline, job name carries a date tag such as 'YYYYMMDD-<domain>-evol-r<N>'). Before the campaign starts, review with the user once per AGENTS.md: the full train task list, the failure-selection criterion, the number of rounds, the fixed model and hyperparameters, and a directly executable command template; after user confirmation, later rounds proceed automatically under the confirmed rules without asking again each round. Re-review whenever the model, dataset, selection criterion, or hyperparameters change.
Reading results: the reward distribution in 'runs/<job>/result.json'; per-trial details in '<job>/<trial>/{agent,verifier}'. Immediately after the rollout finishes, generate the next round's failed-task list and record the tasks that passed this round and left the campaign.
## Evolution discipline
1. **The lib is the only layer.** Shared mechanism -> bank; specific to one task -> do not write it, accept the residual failure. The one-sentence test: is this failure "shared by a class of tasks" or "specific to this one task" -- shared -> bank; specific -> give up on that task, and never write task-specific content (concrete cell addresses, concrete answers, steps unique to one task, concrete file names) into any shared component, skill, or memory. This is both overfitting prevention and the foundation of the claim that "the harness is generic".
2. **Generalize before adding.** If attribution points to a mechanism the bank already has but that does not quite fit -> generalize and improve the existing component (near-duplicate variants are forbidden); if it truly does not exist -> write a new component (first ask yourself "would this be useful on other tasks"; only if generic does it enter bank + manifest + registry).
3. **Prompt material is also a shared asset.** For behavioral failures (finishing without verification, acting before reading the task, writing the deliverable to the wrong place), prefer editing the corresponding 'skills/<name>/SKILL.md' or appending generic clauses to memory.md (English, no source attribution, merged into the matching topic section); do not write task-specific exhortations.
4. **Anti-bloat discipline.** Near-duplicate proposals are deduplicated by the main agent during consolidation; component survival is adjudicated by evaluator performance -- components with no evidence of benefit for two consecutive rounds are moved out of the manifest (files kept for traceability).
5. **Traces are evidence.** Every change must point back to a trajectory attribution in the notes; components that "feel like they should help" are not allowed.
## Attribution checklist (round >=2, go through in order, record in notes)
1. **Did the harness load?** -> In the trial log, confirm the bank_sha256 in 'student_harness_bank.txt' matches the current bank and that the skills index appears in the system prompt; if not mounted: manifest/registry reconciliation or a loader problem.
2. **Did the student use the evolved tools?** -> Search 'llm_calls.jsonl' for the evolved tools' names; if unused: the tool description is not discoverable or the skills index in the system prompt gives no guidance -- fix descriptions / material rather than adding new tools.
3. **Did the middleware fire?** -> If it fired but did not help: the clauses are not actionable enough -- rewrite abstract principles into verbatim rules; if it did not fire: the hook condition does not match the actual trajectory shape.
4. **Is the student dying on mechanics?** (hangs / malformed tool calls / validation loops / repeatedly retrying a failing tool) -> mechanics problem, go back to Evolution discipline and fix the bank.
5. **Dead-trial classification**: infra (does not count) vs turn wall (non-convergence, attribution 2/3/4) vs verifier failure (real failure -- read the verifier details and distinguish "wrong value / wrong location / missing deliverable / wrong format").
## Rounds and acceptance
- Run 3 rounds in total; end early if the failed-task set empties. Record per round: number of input failed tasks, number of tasks that passed this round and left, number of remaining failed tasks, mean reward (vs the bare baseline on the same task set), component additions/changes, recurring failure patterns, and the composition of abnormal trials (infra deaths vs real failures).
- Low scores in round 1 are expected (the v1 artifact comes from static analysis alone, with no attribution iterations); watch the trend, not the absolute value.
- Final delivery after all rounds: the reward curve, the bank's final-state inventory (manifest + component catalog), a map of failure patterns, open issues (including the list of abandoned task-specific residual failures), and suggested material for rewriting into a teacher-side harness.A.2 Evolved components
Table 5 lists every harness component that survives in the final h⋆h^\starh⋆ of each domain. SpreadsheetBench and USPTO keep their accumulated failure notes in an enabled memory component, whereas AppWorld disables the memory slot and injects the same material through the
bootstrap_instruction middleware, which is why the table lists no memory row for it.B. Harness adaptation
The adaptation from h⋆h^\starh⋆ to K\mathcal KK preserves what each component does while changing how and where it acts. Table 6 summarizes the component-level mapping. Beyond individual components, the adaptation can also produce a domain prompt for the harnessing agent that states the domain's review policy.
B.1 Reference harness inventory
K\mathcal KK is mounted as a read-only
/components directory in the harnessing agent's own workspace, which is separate from the student sandbox. Its entry point is an index.json that lists every readable component with an identifier, a kind, a relative path, and a one-line summary. The harnessing agent reads the index before its first decision and then opens individual component files as they become relevant. Table 7 gives the full component inventory for each domain.Review middleware is not reachable as files. These components run automatically on every proposal and, when a condition matches, append a short block of triggered guidance to the update the harnessing agent receives. The domain prompt is appended directly to the harnessing agent's system prompt.
The adaptation is not one-to-one. For example, a student-side guard that probes the sandbox becomes a review-time condition on the visible trajectory, plus an action recipe whenever the correction itself must run in the environment. One guard can also split into several checks. USPTO's
answer_guard is translated into separate conditions for a missing answer file, a hand-written file that bypasses canonicalization, and a commitment made without candidate enumeration.B.2 Worked example: preserving pre-filled cells
Many SpreadsheetBench tasks place a few already-filled example cells inside the requested answer range. Those cells are the specification. The student should infer the filling rule from them, match their format, and write only into the remaining blanks. Overwriting an example is a common failure. The evolved harness therefore inspects the sandbox at finish time, diffs the input and output workbooks inside the answer range, and rejects the finish if any pre-filled literal has changed. Algorithm 2 shows this control flow.
Algorithm 2: Abridged student-side middleware in .
async def on_proposed_finish(state):
input_path, answer_range, output_path = parse_task(state)
result = await run_in_student_sandbox(
PREFILLED_DIFF, input_path, output_path, answer_range
)
if result.violations:
reject_finish(
cells=result.violations,
instruction="Restore the examples and re-derive the rule."
)The harnessing agent cannot run this check privately because it has no separate interface to the student environment, and it should not have one. Otherwise the behavior patterns behind this middleware cannot be internalized. The adapted review middleware instead looks for evidence of the comparison in the student-visible trajectory. If the student tries to finish without that evidence, the middleware adds a private instruction asking the harnessing agent to replace the finish with a verification action:
Algorithm 3: Abridged review middleware in .
def review(candidate, visible_trajectory):
if not candidate.is_finish():
return None
evidence = scan_student_actions(visible_trajectory)
if not evidence.has_prefilled_comparison:
return PrivateInstruction(
decision="REPLACE",
action="Run the pre-filled-cell comparison before finishing."
)The corresponding tool component in K\mathcal KK takes the form of an action recipe. It contains the logic needed to generate a utility script inside the accepted student response. A shortened version is shown in Algorithm 4. The response creates and invokes the script through hhh; its output then becomes visible to the student on the next turn.
Algorithm 4: Student-native verification action generated from the action recipe in .
cat > .agent-tools/prefilled_diff.py <<'PY'
import sys
import openpyxl
from openpyxl.utils import range_boundaries
input_path, output_path, cell_range = sys.argv[1:4]
wb_in = openpyxl.load_workbook(input_path)
wb_out = openpyxl.load_workbook(output_path)
ws_in, ws_out = wb_in.active, wb_out.active
min_c, min_r, max_c, max_r = range_boundaries(cell_range)
changed = []
for row in range(min_r, max_r + 1):
for col in range(min_c, max_c + 1):
before = ws_in.cell(row=row, column=col)
after = ws_out.cell(row=row, column=col)
if before.value is not None and before.data_type != "f":
if after.value != before.value:
changed.append(before.coordinate)
print("PASS" if not changed else f"FAIL changed cells: changed")
raise SystemExit(bool(changed))
PY
python3 .agent-tools/prefilled_diff.py INPUT.xlsx OUTPUT.xlsx B3:B40B.3 Adapting skills and memory
Skills and memory are adapted by changing their audience. The student-side skill states the workflow as direct instructions. Its adaptation in K\mathcal KK tells the harnessing agent how to recognize a missing step and realize that step as a replacement. The following condensed excerpts illustrate the change:
Algorithm 5: Condensed skill adaptation for SpreadsheetBench.
Student-side skill:
Read pre-filled examples before editing.
Infer the rule from them and preserve their values.
Before finishing, compare the output with the input.
Reference-harness guidance:
Check whether the visible trajectory inspected the examples.
If not, replace the current proposal with a bounded inspection.
Before accepting a finish, require an input-output comparison.Memory is adapted in the same way, but starts from failures observed across training tasks. For example, the student-side memory records that agents often overwrite worked examples or verify formulas only against their own outputs. In K\mathcal KK, these observations become review-time patterns with concrete interventions:
Algorithm 6: Condensed memory adaptation for SpreadsheetBench.
Observed failure:
The student overwrites pre-filled examples with its inferred rule.
Review condition:
An edit spans the answer range without preserving existing literals.
Intervention:
Replace with inspection or restore-and-diff actions.
Observed failure:
Verification only rereads values produced by the same script.
Review condition:
No assertion uses worked examples or independent recomputation.
Intervention:
Replace with an assertion-based verification action.Across these adaptations, K\mathcal KK retains the domain-specific condition and corrective behavior. The harnessing agent supplies the final response, and hhh remains the only interface through which that response acts on the environment.
C. Review protocol and harnessing-agent prompt
This section details the interface between the student and the harnessing agent, and gives the system prompt used during training-data collection.
C.1 Session structure
One harnessing-agent session covers one complete student trial. The session is incremental: the harnessing agent keeps its full history across reviews, so earlier updates, its own earlier decisions, and any component files it has read remain in context.
Each review begins with a student update message containing two parts. The first is the set of student-visible events added since the previous review, which on a typical turn is the observation produced by executing the previous accepted response. The second is the current unexecuted proposal, rendered as its reasoning, its visible content, and its tool call, or marked as a final response when it carries no tool call. The first update of a session additionally carries the student system prompt and the task instruction.
In addition, if the student emitted a structurally invalid response, such as more than one tool call (invalid for mini-SWE-agent) or a response with neither content nor a tool call, the update names the defect and states that passing it would preserve the defect. If a review middleware from K\mathcal KK matches the current proposal, the update carries a triggered-guidance block with the middleware's evidence and hint.
C.2 Submission format
The harnessing agent may read adapted harness components under
/components as many times as it needs before deciding, but it must submit exactly one decision per update by calling submit_review. The submission has four fields the agent controls:decision, either PASS{\mathchoice{\text{P{\scriptsize ASS}}}{\text{P{\scriptsize ASS}}}{\text{P{\scriptscriptstyle ASS}}}{\text{PASS}}}PASS or REPLACE{\mathchoice{\text{R{\scriptsize EPLACE}}}{\text{R{\scriptsize EPLACE}}}{\text{R{\scriptscriptstyle EPLACE}}}{\text{REPLACE}}}REPLACE;replacement, the complete student response to execute instead of the proposal;components_used, the identifiers of the components that informed the decision;reason, a private justification of at most 500 characters.
A replacement is a structured object with three fields:
reasoning, content, and an optional tool_call. For mini-SWE-agent, the only admissible tool call is a single execute with a non-empty command, which matches the student's action space under hhh exactly. A replacement with no tool call and non-empty content is a final answer and ends the trial.The schema rejects three malformed harnessing submissions: a PASS{\mathchoice{\text{P{\scriptsize ASS}}}{\text{P{\scriptsize ASS}}}{\text{P{\scriptscriptstyle ASS}}}{\text{PASS}}}PASS that carries a replacement, a REPLACE{\mathchoice{\text{R{\scriptsize EPLACE}}}{\text{R{\scriptsize EPLACE}}}{\text{R{\scriptscriptstyle EPLACE}}}{\text{REPLACE}}}REPLACE that does not, and a replacement with neither visible content nor a tool call. A submission that fails validation is not recorded. When an update produces no valid submission, the runtime re-prompts the harnessing agent for the same candidate, at most three times, and the trial fails if no valid submission is recorded by then. Within one trial, the number of REPLACE{\mathchoice{\text{R{\scriptsize EPLACE}}}{\text{R{\scriptsize EPLACE}}}{\text{R{\scriptscriptstyle EPLACE}}}{\text{REPLACE}}}REPLACE decisions the harnessing agent may issue is a hyperparameter. The budget is stated in the harnessing agent's system prompt, so that it can allocate interventions across the trial. We set it to 5, 3, and 1 to collect trajectories for SpreadsheetBench and AppWorld, and set it to 5 consistently for USPTO.
C.3 System prompt
Algorithm 7 gives the system prompt used to collect training data. Two variants exist. The training-free evaluations in Section 4.2 use a prompt with the same session structure, submission format, and budget, differing only in framing: it describes the goal as turning the components into better student actions rather than as producing a training trajectory, and it omits the paragraph on instructions regarding remaining on-policy. The oracle-answer control in Section 4.4 appends one extra section granting read access to the recorded reference answer for the current task. It also instructs the harnessing agent to use it only to decide whether the student's direction can still converge, to keep every verification step in the replacement, and never to reveal that the answer is known.
Algorithm 7: System prompt of the harnessing agent during training-data collection.
You are a strong harnessing agent supervising a student that solves a task with one `execute` tool and optional subagent delegation through `agent "<task>"` in bash.
The accepted trajectory of this trial will be used directly as SFT training data for the student model. Produce a successful trajectory that stays close to the student's behavior while using the teacher-side harness to introduce reusable task-solving patterns through targeted interventions. The student should do most of the work. When intervention is needed, replace the next response rather than taking over the task, then return control to the student after that response is executed.
On-policy here means that reviews occur along a live student trial and every accepted response uses the student's visible information, native action space, and plausible level of complexity. It does not mean preserving the student's current policy unchanged: a replacement should teach a better behavioral pattern when the harness identifies one that materially improves correctness, progress, recovery, or verification.
## Session input
One teacher session covers one complete student trial. Each `Student update` message contains the student-visible events added since the previous review and the current unexecuted proposal. The first update also contains the student system prompt and task.
The student sandbox is not mounted in your workspace. Treat file contents, command results, installed programs, and service state as known only when they appear in a student-visible observation.
## Components
Read `/components/index.json` before the first decision. Choose and read memory sections or component files as they become relevant during the trial.
Mounted middleware may append `Triggered middleware guidance` to a student update when the current unexecuted proposal matches one of its checks. Use the stated evidence and hint when reviewing that proposal.
Use the components actively as your knowledge of what good behavior looks like. They should affect the accepted trajectory when their guidance is relevant, not merely help you recognize fatal errors. Procedural knowledge from a component may inform a replacement even if the student has not demonstrated it yet, provided the resulting response is a plausible next step in the student's native interface. Component-private facts, paths, review records, and unsupported claims about the current sandbox must never enter the replacement.
## Review
Review every proposed tool call and final answer before it is accepted.
Choose `PASS` when the proposal is a sound next action and no applicable component calls for a meaningful behavioral correction. Pass harmless inefficiencies, stylistic differences, valid alternative methods, and exploratory steps that can produce useful evidence. Also pass recoverable mistakes when observing the result is likely to let the student diagnose and repair them; useful self-recovery is valuable training behavior.
Choose `REPLACE` when a meaningful correction is needed for task success or to instantiate a reusable pattern supplied by the harness. Common reasons include:
- a missed requirement, damaged or skipped deliverable, unsupported conclusion, or premature final answer;
- an unsafe, unbounded, fragile, or budget-wasting command;
- an observed failure that the student ignores or repeats without a useful change;
- a missing inspection, dependency check, test, or verification step needed to ground later work;
- an applicable component identifies a behavior pattern that the proposal violates or omits, and correcting it now would materially improve progress, recoverability, or the value of the trajectory.
You may replace at most `5` student responses during the complete trial. The budget rewards selectivity, not passivity: do not polish sound actions, but do not withhold a useful pattern-level correction merely because the current proposal is not immediately fatal.
When you replace:
- Make the smallest coherent change to the student's next step that installs the correction. This may require a different action, not merely a textual patch.
- Preserve the student's high-level intent, established facts, language, variable names, and command structure when they remain compatible with the correction.
- Stay near the student's demonstrated level, but you may introduce a simple procedure, command, library, or idiom from the components when it is needed to express the target pattern. Make the reasoning understandable from student-visible evidence rather than relying on unexplained teacher expertise.
- Correct the next decision and hand control back. Do not complete several future steps, precompute the task's answer, or replace work the student can perform after seeing the next observation.
- The reasoning must read as the student's own first-person inner monologue, continuing the student's current line of thought -- never as advice, critique, or correction addressed to the student. A natural form is the student catching its own mistake: noticing the constraint, re-reading the evidence, and correcting course.
- A replacement must contain complete student-style reasoning, visible content, and at most one `execute` call, or a final answer without a tool call. Its reasoning, content, and tool call must agree. Base every factual claim on student-visible events.
Do not mention the teacher, review, harness components, component names, private paths, or training in the replacement.
Pass a final answer only after the student-visible observations support every material task requirement. Otherwise replace it with the next inspection, repair, test, or verification action, not with a teacher-written solution that bypasses those steps.
## Submission
Call `submit_review` exactly once for each student update. For `PASS`, set `replacement` to null. For `REPLACE`, provide the complete replacement. Record the components you used and a concise reason in the private submission metadata.D. Trajectory collection and filtering
On SpreadsheetBench and AppWorld, we retain a trial only if its verifier reward is 1.0 and the run finished without an execution exception. USPTO keeps every collected trial, so the ablation in Section 4.4 trains each condition on the same 500 tasks. Training on this full USPTO set also yields higher test performance than keeping successes only.
Every training trajectory must contain only student-visible text. We scan each trajectory for strings that can appear only in the private review and filter out any trajectory that contains the following.
- an absolute
/components/path, which can only refer to the reference harness mount; - the path of the recorded candidate file that stores the unexecuted proposal;
- the name of the review submission tool;
- the review metadata key that records which components a decision used;
- the internal name of the review middleware.
A replacement is written inside the review context, so its reasoning can slip into the reviewer's voice and describe the student's proposal from the outside. We mask the reasoning token span of any accepted response that was produced by a REPLACE{\mathchoice{\text{R{\scriptsize EPLACE}}}{\text{R{\scriptsize EPLACE}}}{\text{R{\scriptscriptstyle EPLACE}}}{\text{REPLACE}}}REPLACE decision and whose reasoning matches a reviewer-perspective pattern. The patterns cover the words proposal and proposed, references to a draft, references to the student's or the candidate's response, reasoning, command, or action, and explicit review verbs applied to a response, such as passing or rewriting it. The resulting training sets contain 487 trajectories for SpreadsheetBench, 282 for AppWorld, and 500 for USPTO, as reported in Section 4.1.
E. Implementation details
E.1 Student harness
The system prompt of mini-SWE-agent hhh is given in Algorithm 8. It is the only instruction the student receives beyond the task itself.
Algorithm 8: System prompt of the student under the target harness .
You are an agent that solves tasks in a Linux sandbox.
You have one tool, `execute`, which runs one bash command. Each call starts a fresh shell, so working-directory and environment changes do not persist between calls.
Work on the task by issuing one `execute` call at a time. When the task is complete, return a short final answer without calling `execute`.
The sandbox also provides `agent "<task>"` for optional subagent delegation.The harness exposes exactly one tool
execute, whose single argument is a non-empty bash command. Each call runs in a fresh shell, so the working directory and environment variables do not carry over; state persists only through the file system. The observation returned to the student is the command's combined output followed by a line giving its exit code. A response with no tool call ends the trial and is taken as the final answer.The sandbox additionally provides an
agent "<task>" CLI utility for single-level subagent delegation. A subagent runs the same base model with the single execute tool, receives a fixed prompt instructing it to return a concise report for its bounded task, and cannot delegate further. At most four subagents may run concurrently per sandbox. This capability ensures that bash plus one-level delegation forms an action space expressive enough to map diverse evolved h⋆h^\starh⋆ components onto hhh. In practice, the evolved h⋆h^\starh⋆ in this work do not rely on subagents, and delegation is used sparingly.E.2 Models and access
We access GPT-5.6 Sol ([53]), DeepSeek-V4-Pro, and DeepSeek-V4-Flash ([54]) through Microsoft Azure. Qwen3.5-9B ([55]) is evaluated as the base model through OpenRouter ([59]), as is Qwen3.6-35B-A3B in the analysis of Appendix F, and the fine-tuned checkpoints are served on NVIDIA A100 GPUs. All harness-evolution runs are performed by Kimi K3 ([56]) using Kimi Code ([25]). We use the default decoding parameters provided by each model provider. Reasoning is enabled for all models, with reasoning effort set to high when applicable.
E.3 Training configuration
We perform LoRA supervised fine-tuning on Qwen3.5-9B with the Tinker supervised-training recipe ([57]), using rank 32, α=32\alpha{=}32α=32, batch size 8, a 65,536-token sequence length, and two epochs. The learning rate follows a linear schedule with 5% warmup to a peak of 2×10−42\times10^{-4}2×10−4, followed by decay to 10−610^{-6}10−6. All runs use seed 42. Each domain is trained separately on its own collected trajectories. Examples longer than the sequence length are dropped.
F. Additional analysis of agent-as-harness
We examine whether the advantage of agent-as-harness extends to weaker models when the same model serves as both the student and the harnessing agent. Table 8 extends the USPTO comparison of Table 1 to two additional, weaker models.
The relative advantage of agent-as-harness over meta-harness ((h,K)−h⋆(h, \mathcal K) - h^\star(h,K)−h⋆) closely correlates with model capability. It remains positive on stronger models but turns negative on weaker ones (+1.0% on GPT-5.6 Sol, +4.0% on DeepSeek-V4-Pro, -1.0% on DeepSeek-V4-Flash, and -12.0% on Qwen3.6-35B-A3B). On the weakest model, the harnessing agent intervenes aggressively, replacing 66.0% of reviewed steps, but its edits are net harmful. Even review without evolved guidance falls below mini-SWE-agent (12.0 vs. 16.0). Effective review and intervention therefore require sufficient underlying model capability.
The current HARNESS-ZERO{\mathchoice{\text{H{\scriptsize ARNESS}-Z{\scriptsize ERO}}}{\text{H{\scriptsize ARNESS}-Z{\scriptsize ERO}}}{\text{H{\scriptscriptstyle ARNESS}-Z{\scriptscriptstyle ERO}}}{\text{HARNESS-ZERO}}}HARNESS-ZERO design reviews every student proposal before execution, requiring at least two model calls per step: one to generate the proposal and another to review it. Each review processes the trajectory prefix and reference harness, and must finish before the next student step, which increases both token use and latency. On USPTO, agent-as-harness with evolved guidance takes 237.2 s per trial on average, about 2.4×2.4\times2.4× the 100.1 s required by mini-SWE-agent. Reducing the overhead, for example by reviewing only a subset of steps, is left to future work.
G. Harness-exclusive behavior patterns
Here we detail how the behavior patterns of Section 4.5 are mined and explain the mined patterns on the three benchmarks. All measurements use the trajectories of three deployments from Table 2: mini-SWE-agent (hhh), meta-harness (h⋆h^\starh⋆), and HARNESS-ZERO{\mathchoice{\text{H{\scriptsize ARNESS}-Z{\scriptsize ERO}}}{\text{H{\scriptsize ARNESS}-Z{\scriptsize ERO}}}{\text{H{\scriptscriptstyle ARNESS}-Z{\scriptscriptstyle ERO}}}{\text{HARNESS-ZERO}}}HARNESS-ZERO (hhh).
G.1 Approach
For each domain we start from the enabled components of the student-side harness h⋆h^\starh⋆ and read every enabled memory file, skill file, tool, and middleware. Then, we translate each rule into a deterministic binary detector over the trajectory (agent messages, bash commands, and observations). Detectors match the target behavior semantically. For example, AppWorld pagination counts if the trajectory uses the
aw.pages helper or an explicit loop over page_index; a USPTO answer write counts if it uses a shell redirect or Python write_text. The USPTO tools smiles_check, propose_retrosynthesis, and verify_answer exist only under h⋆h^\starh⋆. A call to them counts on h⋆h^\starh⋆ trajectories, and an equivalent bash or Python command counts under hhh.For each detector, we define the support tasks as the test tasks on which, for the base model, the pattern appears under h⋆h^\starh⋆ but not under hhh. We keep a detector only if it has at least 10 support tasks. If two detectors give the same result on every applicable task, they are measuring the same behavior, so we merge them into one pattern. This procedure yields 18 patterns on SpreadsheetBench, 6 on USPTO, and 4 on AppWorld. A pattern's recovery rate is the fraction of its support tasks on which HARNESS-ZERO{\mathchoice{\text{H{\scriptsize ARNESS}-Z{\scriptsize ERO}}}{\text{H{\scriptsize ARNESS}-Z{\scriptsize ERO}}}{\text{H{\scriptscriptstyle ARNESS}-Z{\scriptscriptstyle ERO}}}{\text{HARNESS-ZERO}}}HARNESS-ZERO exhibits the behavior. By construction, the base model under hhh and h⋆h^\starh⋆ exhibits it on 0% and 100% of these tasks, respectively.
G.2 Mined patterns
Table 10, Table 11, and Table 9 list every retained pattern.
References
[1] Weng, Lilian (2026). Harness Engineering for Self-Improvement. https://lilianweng.github.io/posts/2026-07-04-harness/. Lil'Log.
[2] Ning et al. (2026). Code as Agent Harness. arXiv preprint arXiv:2605.18747.
[3] Yang et al. (2024). Swe-agent: Agent-computer interfaces enable automated software engineering. Advances in Neural Information Processing Systems. 37. pp. 50528–50652.
[4] Lu et al. (2026). Towards end-to-end automation of AI research. Nature. 651(8107). pp. 914–919.
[5] Ye et al. (2026). Meta Context Engineering via Agentic Skill Evolution. In Proceedings of the 43rd International Conference on Machine Learning. https://arxiv.org/abs/2601.21557.
[6] Ma et al. (2026). LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks. arXiv preprint arXiv:2608.01964. https://arxiv.org/abs/2608.01964.
[7] Karten et al. (2026). Continual Harness: Online Adaptation for Self-Improving Foundation Agents. arXiv preprint arXiv:2605.09998.
[8] Karten et al. (2026). Prime Agent: A Self-Improving RLM Harness. arXiv preprint arXiv:2608.23552.
[9] Yan et al. (2026). Harness-of-Harness: Multi-Day Autonomous Software Development with Continual Improvement. arXiv preprint arXiv:2609.01481. https://arxiv.org/abs/2609.01481.
[10] Ye et al. (2024). Reevo: Large language models as hyper-heuristics with reflective evolution. Advances in neural information processing systems. 37. pp. 43571–43608.
[11] Lee et al. (2026). Meta-Harness: End-to-End Optimization of Model Harnesses. arXiv preprint arXiv:2603.28052.
[12] Lin et al. (2026). Agentic Harness Engineering: Observability-Driven Automatic Evolution of Coding-Agent Harnesses. arXiv preprint arXiv:2604.25850.
[13] Zhang et al. (2026). Self-Harness: Harnesses That Improve Themselves. arXiv preprint arXiv:2606.09498. https://arxiv.org/abs/2606.09498.
[14] Luo et al. (2026). HarnessBank: Semantic Gene-Bank Search with Gated Verification for Agent-Harness Self-Evolution. arXiv preprint arXiv:2607.13683. https://arxiv.org/abs/2607.13683.
[15] Liu, Ming (2026). More Is Not Always Better: Cross-Component Interference in LLM Agent Scaffolding. arXiv preprint arXiv:2605.05716. https://arxiv.org/abs/2605.05716.
[16] Yao et al. (2026). Harness-Bench: Measuring Harness Effects across Models in Realistic Agent Workflows. arXiv preprint arXiv:2605.27922. https://arxiv.org/abs/2605.27922.
[17] Zhang et al. (2026). Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models. In The Fourteenth International Conference on Learning Representations. https://openreview.net/forum?id=eC4ygDs02R.
[18] Lin et al. (2026). Task-Aware Harness Provisioning for LLM Agents in Mission-Critical Infrastructure Operations. arXiv preprint arXiv:2608.17433. https://arxiv.org/abs/2608.17433.
[19] Chang et al. (2026). From LLM Inference to Agentic Workloads: Characterization and Implications for Serving Systems. arXiv preprint arXiv:2608.15127. https://arxiv.org/abs/2608.15127.
[20] Zhang et al. (2025). Cut the Crap: An Economical Communication Pipeline for LLM-based Multi-Agent Systems. In The Thirteenth International Conference on Learning Representations. https://openreview.net/forum?id=LkzuPorQ5L.
[22] Zhou et al. (2026). Externalization in LLM Agents: A Unified Review of Memory, Skills, Protocols and Harness Engineering. arXiv preprint arXiv:2604.08224.
[25] Moonshot AI (2026). Kimi Code CLI. https://github.com/MoonshotAI/kimi-code. Accessed: 2026-09-10.
[26] Earendil Works (2026). Pi: An Agent Harness and Coding Agent. https://github.com/earendil-works/pi. Accessed: 2026-09-10.
[27] Anomaly (2026). OpenCode: The Open Source Coding Agent. https://github.com/anomalyco/opencode. Accessed: 2026-09-10.
[28] Pan et al. (2026). Natural-Language Agent Harnesses. arXiv preprint arXiv:2603.25723. https://arxiv.org/abs/2603.25723.
[29] Khattab et al. (2023). Dspy: Compiling declarative language model calls into self-improving pipelines. arXiv preprint arXiv:2310.03714.
[30] Fernando et al. (2023). Promptbreeder: Self-Referential Self-Improvement via Prompt Evolution. arXiv preprint arXiv:2309.16797.
[31] Agrawal et al. (2025). Gepa: Reflective prompt evolution can outperform reinforcement learning. arXiv preprint arXiv:2507.19457.
[32] Hu et al. (2025). Automated Design of Agentic Systems. In International Conference on Learning Representations.
[33] Zhang et al. (2025). AFlow: Automating Agentic Workflow Generation. In The Thirteenth International Conference on Learning Representations. https://openreview.net/forum?id=z5uVAKwmjf.
[34] Zhang et al. (2026). JIT-Agent: Scaling Harness Intelligence via Just-in-Time Harness Evolution. arXiv preprint arXiv:2608.25593. https://arxiv.org/abs/2608.25593.
[35] Qian et al. (2026). AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses. arXiv preprint arXiv:2608.12307.
[36] Hebbar et al. (2026). SIA: Self Improving AI with Harness & Weight Updates. arXiv preprint arXiv:2605.27276.
[37] Chen et al. (2026). Co-Harness: Co-Evolving Harnesses and Model Weights for LLM Agents. arXiv preprint arXiv:2607.22688.
[38] Kim et al. (2026). WHALE: A Simple Recipe for Joint Harness–Weight Optimization. arXiv preprint arXiv:2609.00196. https://arxiv.org/abs/2609.00196.
[39] Chen et al. (2026). HarnessForge: Joint Harness and Policy Evolution for Adaptive Agent Systems. arXiv preprint arXiv:2606.01779. https://arxiv.org/abs/2606.01779.
[40] Chen et al. (2026). HarnessX: A Composable, Adaptive, and Evolvable Agent Harness Foundry. arXiv preprint arXiv:2606.14249. https://arxiv.org/abs/2606.14249.
[41] Luo et al. (2026). Harness-Aware Self-Evolving: Co-Evolving Model Weights, Harness, and Task Solutions. arXiv preprint arXiv:2607.03935. https://arxiv.org/abs/2607.03935.
[42] Mao et al. (2026). SafeEvolve: Harness-Policy Co-Evolution from Agent Experience for Safety Alignment. arXiv preprint arXiv:2609.02786. https://arxiv.org/abs/2609.02786.
[43] Le et al. (2026). What Does Multi-Harness RL Learn? Credit Assignment and Portability in Coding Agents. arXiv preprint arXiv:2609.04518. https://arxiv.org/abs/2609.04518.
[44] Ning et al. (2026). EvoHarness-RL: Learning Self-Evolving Runtime Harness for Long-Horizon LLM Agents. arXiv preprint arXiv:2608.05446.
[45] Zhao et al. (2026). Training with Harnesses: On-Policy Harness Self-Distillation for Complex Reasoning. arXiv preprint arXiv:2605.08741.
[46] LangChain (2025). Deep Agents. https://github.com/langchain-ai/deepagents. Accessed: 2026-01-19.
[47] Ma, Zeyao and others (2024). SpreadsheetBench: Towards Challenging Real World Spreadsheet Manipulation. arXiv:2406.14991.
[48] Trivedi et al. (2024). AppWorld: A Controllable World of Apps and People for Benchmarking Interactive Coding Agents. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics. pp. 16022–16076.
[49] Lowe, Daniel Mark (2012). Extraction of chemical structures and reactions from the literature.
[50] Jin et al. (2017). Predicting Organic Reaction Outcomes with Weisfeiler-Lehman Network. In Advances in Neural Information Processing Systems. pp. 2607–2616.
[51] Harbor Framework Team (2026). Harbor: A Framework for Evaluating and Optimizing Agents and Models in Container Environments. https://github.com/harbor-framework/harbor. doi:10.5281/zenodo.20953922.
[52] Shi et al. (2026). Harbor Adapters and Harbor-Index: Infrastructure and a Curated Meta-Dataset for Large-Scale Agentic Evaluation. arXiv preprint arXiv:2609.04298.
[53] OpenAI (2026). GPT-5.6 System Card. https://deploymentsafety.openai.com/gpt-5-6. Accessed: 2026-09-14.
[54] DeepSeek-AI (2026). DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence. https://arxiv.org/abs/2606.19348. arXiv:2606.19348.
[55] Qwen Team (2026). Qwen3.5: Towards Native Multimodal Agents. https://qwen.ai/blog?id=qwen3.5. Accessed: 2026-09-14.
[56] Kimi Team (2026). Kimi K3: Open Frontier Intelligence. https://arxiv.org/abs/2607.24653. arXiv:2607.24653.
[57] Thinking Machines Lab (2025). Announcing Tinker. https://thinkingmachines.ai/news/announcing-tinker/. Accessed: 2026-09-17.
[58] Zuo et al. (2026). Qwen-AgentWorld: Language World Models for General Agents. arXiv preprint arXiv:2606.24597. https://arxiv.org/abs/2606.24597.
[59] OpenRouter, Inc. (2025). OpenRouter: The Unified Interface for LLMs. https://openrouter.ai/. Accessed: 2026-01-16.











