Toward Generalist Autonomous Research via Hypothesis-Tree Refinement
Jiajie Jin1,†,‡^{1,\dagger,\ddagger}1,†,‡, Yuyang Hu1,†^{1,\dagger}1,†, Kai Qiu2^{2}2, Qi Dai2^{2}2, Chong Luo2^{2}2, Guanting Dong1^{1}1, Xiaoxi Li1^{1}1, Tong Zhao1^{1}1, Xiaolong Ma2^{2}2, Gongrui Zhang2^{2}2, Zhirong Wu2^{2}2, Bei Liu2^{2}2, Zhengyuan Yang2^{2}2, Linjie Li2^{2}2, Lijuan Wang2^{2}2, Hongjin Qian1^{1}1, Yutao Zhu1^{1}1, Zhicheng Dou1,∗^{1,*}1,∗
1^{1}1 Gaoling School of Artificial Intelligence, Renmin University of China
2^{2}2 Microsoft Research
1^{1}1 Gaoling School of Artificial Intelligence, Renmin University of China
2^{2}2 Microsoft Research
†^{\dagger}† Equal contribution
‡^{\ddagger}‡ Work done during an internship at MSRA
∗^{*}∗ Corresponding author
‡^{\ddagger}‡ Work done during an internship at MSRA
∗^{*}∗ Corresponding author
Abstract
Scientific progress depends on a repeated loop of exploration, experimentation, and abstraction. Researchers test candidate directions, interpret the evidence, and carry the resulting lessons into later attempts. We study how an AI agent can run this loop autonomously over long horizons. We introduce Arbor, a general framework for autonomous research that combines a long-lived coordinator, short-lived executors, and Hypothesis Tree Refinement (HTR), a persistent tree that links hypotheses, artifacts, evidence, and distilled insights across time. The coordinator manages global research strategy over the tree, while executors implement and test individual hypotheses in isolated worktrees. As results return, Arbor updates the tree, propagates reusable lessons, refines the search frontier, and admits verified improvements. This design turns autonomous research from a sequence of local attempts into a cumulative process in which strategy, execution, and evidence are carried across time. We evaluate Arbor under Autonomous Optimization (AO), an operational setting where an agent improves an initial research artifact through iterative experimentation without step-level human supervision. Across six real research tasks in model training, harness engineering, and data synthesis, Arbor achieves the best held-out result on all six tasks, attaining more than 2.5×2.5\times2.5× the average relative held-out gain of Codex and Claude Code under the same task interface and resource budget. On MLE-Bench Lite, Arbor reaches 86.36% Any Medal with GPT-5.5, the strongest result in our comparison.
1. Introduction
Scientific research is a central form of long-horizon human intelligence ([1, 2]). Its difficulty lies not only in solving isolated problems, but in sustaining progress across uncertain hypotheses, costly experiments, failed attempts, and delayed feedback. A researcher must maintain an evolving understanding of the problem so that each attempt can reshape what should be tried next ([3]). Recent LLM agents can now edit code, call tools, retrieve information, and run experiments for extended periods ([4, 5, 6]), and systems such as Codex ([7]), Claude Code ([8]), and OpenHands ([9]) make sustained progress in real codebases, making autonomous research an increasingly concrete systems problem. Yet longer execution alone does not guarantee research progress. The open challenge is how an agent can maintain a research state that turns many local attempts into cumulative hypothesis refinement and verified artifact improvement.
We formalize this problem as Autonomous Optimization (AO), which captures the core operational form of autonomous research. In AO, an agent begins with an initial artifact and a research objective, then improves the artifact through experimental feedback without step-level human supervision. This setting is difficult because research feedback is delayed, experiments can be expensive, and failed attempts often contain information that should guide later search. As the horizon grows, an agent that treats each trial as an independent local attempt loses the structure of the research process. Effective AO therefore requires a persistent research state that records what has been tried, what evidence was obtained, and how each result changes the space of future hypotheses.
Despite recent progress, current agent systems still do not provide a general framework for running autonomous research over long horizons ([10, 11]). General coding agents can edit code, invoke tools, and run experiments for many hours, but their autonomy is mostly expressed as persistent task execution. Scientific-agent systems move closer to research automation, yet many still follow predefined workflows or revise a single line of work at a time ([12, 13, 14]). They therefore lack the mechanism that makes human research cumulative: the ability to maintain competing directions, test them through concrete experiments ([15, 16]), interpret both successes and failures ([3]), and let those lessons reshape later exploration. For AO, the key challenge is to build this mechanism into the agent system itself, so that long-running experimentation becomes a self-directed research process rather than an extended sequence of local attempts.
We argue that a general AO system should automate the long-horizon work that a human researcher normally performs during iterative research. Starting from an open objective, it should form research directions, test them through concrete artifact changes, and turn the resulting evidence into memory that shapes later exploration. Progress should not depend on a human repeatedly choosing the next attempt or interpreting what previous trials mean. Instead, the system needs a framework that keeps directions, experiments, artifacts, results, and failures connected across time, turning autonomous research into a persistent cycle of exploration and verified improvement.
We introduce Arbor, a general framework and open-source research system for AO.
Arbor separates autonomous research into a long-lived coordinator and short-lived executors. The coordinator owns the global research state and decides how the search frontier should evolve, while each executor tests one hypothesis in an isolated worktree and returns structured evidence. Arbor makes this two-level process cumulative through Hypothesis Tree Refinement (HTR). HTR represents the research process as a persistent tree in which each node binds a hypothesis, the artifact version that realizes it, the experimental evidence it produces, and the distilled insight that should shape later decisions. When executor results return, Arbor writes evidence back to the executed nodes, abstracts local findings upward, and uses the updated tree to decide which directions to expand, prune, or merge, promoting a candidate to the current best only when it improves a held-out evaluation. The tree therefore acts as the operational research state of the system: it is simultaneously the search frontier, the memory of past attempts, and the audit trail for verified artifact improvement.To evaluate
Arbor, we construct six AO tasks from real research settings across model training ([17, 18]), harness engineering ([19, 20]), and data synthesis. Each task specifies an initial artifact, a natural-language objective, a task-native metric, and a development/test protocol that separates exploratory feedback from final scoring. Arbor achieves the best held-out result on all six tasks, with more than 2.5×2.5\times2.5× the average relative gain of Codex and Claude Code under the same task interface and resource budget. On MLE-Bench Lite ([21]), Arbor further reaches 86.36% Any Medal with GPT-5.5, the strongest result in our comparison. Ablations, backbone studies, transfer experiments, and cost analyses show that these gains come from Arbor's evidence-structured research process: hypotheses remain grounded in executable artifacts, local findings become reusable insights, and later decisions are made over an explicit research state.Our contributions are summarized as follows:
- We formulate Autonomous Optimization (AO) as a class of long-horizon research tasks in which an agent must iteratively improve an artifact under a fixed objective and evaluator without step-level human supervision.
- We introduce Arbor, a general framework for AO that organizes research through Hypothesis Tree Refinement (HTR), pairing a persistent coordinator with isolated executors so that hypotheses, artifact versions, experimental evidence, and distilled insights accumulate into an auditable research state; we release it as an open-source research system.
- We construct six AO tasks from real research settings and show, together with MLE-Bench Lite, that
Arbordelivers the strongest held-out gains and that persistent hypothesis management and insight propagation are the key drivers of its performance.
2. Related Work
2.1 Autonomous Research Agent
LLM-based automated research systems first appeared as end-to-end pipelines. The AI Scientist ([12]) connected idea generation, implementation, execution, result interpretation and paper writing in a mostly automated loop, while Agent Laboratory ([13]) organized similar stages as a human-supervised multi-agent research-assistant workflow. The next wave made the search process more explicit: AIDE ([22]) explored ML engineering as iterative code search, AI Scientist-v2 ([23]) introduced agentic tree search over research plans and experiments, and multi-agent systems such as AI-Researcher, R&D-Agent ([24]) and Loongflow ([25]) refined the literature-to-experiment loop through more specialized roles and summarization mechanisms. In parallel, FunSearch ([26]) and AlphaEvolve ([27]) treated LLMs as program mutation operators selected by executable fitness signals, and SciMaster ([28]) broadened automated research from ML experimentation to general-purpose scientific reasoning with tool-augmented, breadth-and-depth search.
Recent systems expand the object of search itself. MARS ([29]) modularizes automated AI research into reflective components; AutoHarness ([11]), Meta-Harness ([10]) and AHE ([30]) search or evolve the code harness surrounding an agent; and DataMaster ([31]) moves the search target to data, using DataTree, Data Pool and Global Memory to organize autonomous data discovery and validation. Arbor instead stores research state in a persistent hypothesis tree. A long-running coordinator expands and updates the tree, while short-lived executors implement individual hypotheses in isolated git worktrees. This design makes hypotheses, failures, evidence and merge decisions auditable. It also addresses gaps noted by recent surveys and sandbagging studies: weak evidence preservation, loose dev/test discipline and silent metric chasing ([2, 32]).
2.2 Long-Horizon Agent
As language agents have improved, the key question has shifted from whether they can complete isolated tool-use episodes to how long they can remain coherent on real tasks ([33, 34, 35]). Early systems such as Reflexion ([36]) and Generative Agents ([37]) extended single-run behavior with natural-language memories or reflections across trials, making experience accumulation part of the agent loop. Later human-calibrated evaluations made this horizon measurable: some tasks compare agents with human time-to-complete ([38, 39, 40]), while others ([41, 42]) show that agents still struggle to preserve and reuse evidence across long optimization histories, even when the environment provides executable feedback.
Recent approaches therefore increasingly treat long-horizon agency as a problem of externalized state organization rather than only prompt design. Some systems organize prior experience into persistent context, using curated playbooks ([43]), state-adaptive trajectory retrieval ([44]), or cognitive caches ([45]) so that later decisions can draw on earlier failures and successes. Others push the same idea into the scaffold around the model, coordinating agents through persistent workspaces ([1]), recursively modifying agent code ([46]), or evolving harnesses that shape tool use and execution constraints ([11]). Arbor follows this state-externalization trend, but its persistent object is specifically a research tree: each node binds a hypothesis, implementation branch, result, score, related work and learned insight, so progress accumulates through branch expansion, insight backpropagation, merge decisions and pruning rather than through an ever-growing context window.
2.3 Benchmark for Autonomous Research
Research-agent benchmarks have progressed along several complementary directions. MLAgentBench ([47]), MLE-bench ([21]) and MLE-Dojo ([48]) evaluate agents on executable ML-engineering workflows with objective task metrics. ScienceAgentBench ([49]), PaperBench ([50]), and FrontierScience ([51])focus on programmatic discovery, or paper reproduction under expert-designed questions and rubrics. RE-Bench ([38]), HCAST ([39]), AlgoTune ([52]), NanoGPT speedrunning ([41]), PostTrainBench ([42]) and Frontier-Eng ([53]) further emphasize long-horizon engineering, human-calibrated difficulty, executable feedback and iterative self-improvement.
These benchmarks measure important outcomes, but many evaluations still use incomplete settings: some lack a clear dev/test split, making iterative search prone to overfitting, while others rely on one or two task types and leave generality under-tested. Arbor therefore evaluates across multiple tasks and enforces a stricter protocol: data and evaluation harnesses are immutable, the dev set supports iterative search, the test set is reserved for merge or final validation, and each experiment is tied to branch-level artifacts and hypothesis-tree records.
3. Task Formulation
We model auto-research as an instance of Autonomous Optimization (AO), which can be represented as a tuple
The material M0\mathcal{M}_0M0 is the mutable artifact the agent may inspect and modify, typically a codebase together with its associated data. The objective O\mathcal{O}O specifies what it means for a modified material M′\mathcal{M}'M′ to be better, for example a metric direction defined over the artifact's output. The two evaluators instantiate the same objective on different evidence: Edev\mathcal{E}_{\mathrm{dev}}Edev returns feedback the agent may freely use during search, while the held-out Etest\mathcal{E}_{\mathrm{test}}Etest measures whether the dev-driven improvement transfers beyond the feedback used for exploration. Let Sdev(M′)S_{\mathrm{dev}}(\mathcal{M}')Sdev(M′) and Stest(M′)S_{\mathrm{test}}(\mathcal{M}')Stest(M′) denote the scalar scores returned by these two evaluators under the metric direction specified by O\mathcal{O}O, so that larger values are better; a candidate that exploits idiosyncrasies of the dev split may improve SdevS_{\mathrm{dev}}Sdev but is not a successful AO solution unless the gain also transfers to StestS_{\mathrm{test}}Stest.
During a run, the agent adaptively generates, implements, and evaluates candidate materials using Edev\mathcal{E}_{\mathrm{dev}}Edev. Let A\mathcal{A}A be the set of candidates produced. The artifact-level goal is to return
subject to the constraint that hypotheses and implementation decisions are made without using Etest\mathcal{E}_{\mathrm{test}}Etest as an exploration oracle.
4. The Arbor Framework
4.1 Overview
We propose
Arbor, a general framework for autonomous research under the AO interface defined in Section 3. AO differs from ordinary agentic tool use in that the target is not a single response or code patch, but a sustained research trajectory. An agent must propose hypotheses, materialize them as artifact changes, interpret experimental feedback, and decide which directions should be refined, merged, or abandoned. The central design problem is therefore how to convert many transient trials into cumulative research progress.This problem imposes three requirements on the system design:
- Branching with coherence. Research exploration must branch because multiple competing hypotheses may be plausible at the same time. However, unrestricted branching can degenerate into an unstructured log of attempts. The system must therefore maintain a frontier in which competing directions coexist while remaining organized, comparable, and actionable.
- Global strategy with local execution. Strategic decisions depend on evidence accumulated across the whole run, whereas implementing a single hypothesis requires short-horizon code editing, debugging, and evaluation. These two levels should be separated so that low-level execution traces do not obscure the global research state, and experimental outcomes remain attributable to the hypotheses that produced them.
- Exploration with held-out admission. Development feedback should guide hypothesis search, but artifact-level progress should be admitted only when it transfers beyond the feedback used during exploration. The system must therefore distinguish exploratory improvement on EdevE_{\mathrm{dev}}Edev from verified improvement under the held-out evaluator EtestE_{\mathrm{test}}Etest.
Arbor addresses these requirements through Hypothesis Tree Refinement (HTR), as illustrated in Figure 2. Its central state is a persistent hypothesis tree whose nodes bind together a research hypothesis, the artifact version that realizes it, the evaluation evidence it produces, and the distilled insight that should influence later decisions. A long-lived coordinator maintains this tree as the global research state: it observes the current frontier, proposes refinements, selects promising leaves, integrates returned evidence, propagates insights upward, and decides whether to continue, prune, or merge a candidate branch. Short-lived executors test selected hypotheses in isolated worktrees and return compact reports containing scores, factual results, distilled insights, and artifact references. A held-out merge gate promotes a candidate to the current best artifact only when its improvement transfers beyond the development evaluator. In this way, Arbor organizes AO as evidence-structured refinement over a durable research state rather than repeated local trial-and-error. Section 4.2 describes the hypothesis-tree representation, and Section 4.3 presents the coordinator–executor loop that maintains it.4.2 Hypothesis Tree as Research State
The center of Figure 2 shows the hypothesis tree that serves as
Arbor's persistent research state. In AO, the intermediate state is not only the latest artifact or its evaluation score, but also the structure of exploration: which hypotheses have been considered, how they relate to one another, what evidence they produced, and what lessons should constrain future trials. A tree is a natural representation for this state because it preserves both the branching structure of research exploration and the abstraction hierarchy from broad directions to executable interventions.Let T=(V,E)\mathcal{T}=(\mathcal{V}, \mathcal{E})T=(V,E) denote a rooted hypothesis tree with root node n0n_0n0. Each node n∈Vn\in\mathcal{V}n∈V is a research unit
where the three fields separate the semantic content of a hypothesis, the reusable evidence derived from it, and the executable record that grounds it:
- Hypothesis hnh_nhn. The hypothesis describes a verifiable or falsifiable claim about how the material should be changed to improve the objective. Its granularity depends on the node depth: nodes close to the root describe broad research directions, while deeper nodes specify concrete interventions that can be implemented and evaluated by an executor. This allows
Arborto organize exploration as progressive refinement rather than as a flat sequence of independent trials. - Insight ιn\iota_nιn. The insight stores the reusable interpretation of evidence associated with the hypothesis. For an executed leaf, it summarizes what was tried, what happened, and why the result supports, weakens, or constrains the hypothesis. For an internal node, it abstracts over the insights of its children and summarizes the current understanding of that research direction. Thus, ιn\iota_nιn is not an execution transcript, but a compact semantic memory for later hypothesis generation and selection.
- Metadata μn\mu_nμn. The metadata connects the semantic hypothesis to executable evidence. It includes the node status, development score when available, factual result record, implementation reference such as a git branch or commit, and optional background evidence. The material itself is not duplicated in the tree; instead, the tree stores references to external artifact states produced in isolated worktrees. This keeps the research state compact while ensuring that each hypothesis remains grounded in a verifiable implementation.
The tree separates internal direction nodes from executable leaf nodes. Internal nodes maintain abstract research directions and accumulated lessons, whereas leaves represent candidate interventions that can be dispatched for implementation and evaluation. After a leaf is executed, its score, result, artifact reference, and distilled insight are written back to the corresponding node. The insight is then propagated upward by updating the ancestors along the path to the root. Through this abstraction process, local experimental outcomes become direction-level lessons and eventually contribute to a compact global understanding of the run.
In this way, the hypothesis tree serves three roles simultaneously. It is a search frontier that records which directions remain active, validated, or pruned; a long-term memory that stores reusable evidence from both successes and failures; and an auditable research record that links each artifact change to the hypothesis and evidence that motivated it. This persistent state provides the substrate on which the coordinator can make strategic decisions across long-horizon autonomous optimization.
4.3 Hypothesis Tree Refinement
To maintain the tree over a long-horizon AO run,
Arbor separates global frontier control from local experimental execution. A persistent coordinator owns the shared tree and decides where to expand, which evidence to trust, which directions to prune, and when a candidate should be merged. Short-lived executors are invoked only to test individual hypotheses: each executor receives one tree node, materializes the corresponding intervention in an isolated git worktree, evaluates it, and returns structured evidence to the coordinator.During the research process, the coordinator sees the whole research frontier but does not directly perform every low-level implementation step; the executor performs grounded engineering work but does not modify the shared tree or redirect the search objective. As a result, exploratory code changes remain isolated until they pass the merge gate, while the tree records only decision-relevant evidence: scores, factual outcomes, artifact references, and distilled insights. This boundary allows
Arbor to turn transient execution traces into a persistent research state without reducing the tree to a raw log of tool calls.4.3.1 Coordinator: Evidence-Aware Frontier Control
The coordinator updates T ree\mathcal{T}\!\textit{ree}Tree through a repeated six-step procedure: OBSERVE{\mathchoice{\text{O{\scriptsize BSERVE}}}{\text{O{\scriptsize BSERVE}}}{\text{O{\scriptscriptstyle BSERVE}}}{\text{OBSERVE}}}OBSERVE, IDEATE{\mathchoice{\text{I{\scriptsize DEATE}}}{\text{I{\scriptsize DEATE}}}{\text{I{\scriptscriptstyle DEATE}}}{\text{IDEATE}}}IDEATE, SELECT{\mathchoice{\text{S{\scriptsize ELECT}}}{\text{S{\scriptsize ELECT}}}{\text{S{\scriptscriptstyle ELECT}}}{\text{SELECT}}}SELECT, DISPATCH{\mathchoice{\text{D{\scriptsize ISPATCH}}}{\text{D{\scriptsize ISPATCH}}}{\text{D{\scriptscriptstyle ISPATCH}}}{\text{DISPATCH}}}DISPATCH, BACKPROPAGATE{\mathchoice{\text{B{\scriptsize ACKPROPAGATE}}}{\text{B{\scriptsize ACKPROPAGATE}}}{\text{B{\scriptscriptstyle ACKPROPAGATE}}}{\text{BACKPROPAGATE}}}BACKPROPAGATE, and DECIDE{\mathchoice{\text{D{\scriptsize ECIDE}}}{\text{D{\scriptsize ECIDE}}}{\text{D{\scriptscriptstyle ECIDE}}}{\text{DECIDE}}}DECIDE. Each step operates on the tree through a narrow interface for adding nodes, dispatching executors, updating node evidence, propagating insights, pruning subtrees, and merging verified branches. The key point is that the LLM policy chooses how to interpret the research state, while all durable state changes are expressed as controlled mutations of the hypothesis tree.
OBSERVE{\mathchoice{\text{O{\scriptsize BSERVE}}}{\text{O{\scriptsize BSERVE}}}{\text{O{\scriptscriptstyle BSERVE}}}{\text{OBSERVE}}}OBSERVE. At the beginning of each cycle, the coordinator re-grounds itself in the current research state by reading a structured projection of T ree\mathcal{T}\!\textit{ree}Tree, including active frontier nodes, recently returned evidence, ancestor insights, and the current best artifact Mbest\mathcal{M}_{\mathrm{best}}Mbest. This step makes the tree the authoritative state after context compression and prevents the coordinator from relying on a lossy conversational history.
IDEATE{\mathchoice{\text{I{\scriptsize DEATE}}}{\text{I{\scriptsize DEATE}}}{\text{I{\scriptscriptstyle DEATE}}}{\text{IDEATE}}}IDEATE. The coordinator selects a parent node and proposes a small set of child hypotheses beneath it. Each child represents a refinement, alternative, or correction of the parent hypothesis and is initialized as a pending node. Unlike free-form brainstorming, ideation is conditioned on accumulated tree evidence: validated insights provide assumptions to build on, pruned nodes provide negative constraints, and recent executor reports suggest which interventions are feasible or under-tested.
SELECT{\mathchoice{\text{S{\scriptsize ELECT}}}{\text{S{\scriptsize ELECT}}}{\text{S{\scriptscriptstyle ELECT}}}{\text{SELECT}}}SELECT. The coordinator chooses pending nodes to execute next. Selection balances the expected utility of a hypothesis with the evidence already accumulated around its ancestors and siblings. A direction may be selected because it has strong prior evidence, because its siblings expose an unresolved ambiguity, or because its failure would clarify an important assumption. Thus selection is not merely score maximization; it is frontier control under partial and delayed feedback.
DISPATCH{\mathchoice{\text{D{\scriptsize ISPATCH}}}{\text{D{\scriptsize ISPATCH}}}{\text{D{\scriptscriptstyle ISPATCH}}}{\text{DISPATCH}}}DISPATCH. Selected hypotheses are dispatched to independent executors. Each executor materializes its assigned hypothesis in a fresh worktree, evaluates the modified artifact on Edev\mathcal{E}_{\mathrm{dev}}Edev, and returns a compact report containing the dev score, factual result, distilled insight, and branch reference. Parallel execution of sibling hypotheses provides comparative evidence within the same research direction, which is useful for later pruning and abstraction.
BACKPROPAGATE{\mathchoice{\text{B{\scriptsize ACKPROPAGATE}}}{\text{B{\scriptsize ACKPROPAGATE}}}{\text{B{\scriptscriptstyle ACKPROPAGATE}}}{\text{BACKPROPAGATE}}}BACKPROPAGATE. When executor reports return, the coordinator writes their evidence into the corresponding leaf nodes and updates insights along the path to the root. The propagated signal is not only a scalar score. It also includes causal attributions, applicability conditions, and reusable lessons extracted from the experiment. A leaf-level observation such as a data-interface mismatch can therefore become a direction-level constraint, and eventually a global prior that shapes future ideation.
DECIDE{\mathchoice{\text{D{\scriptsize ECIDE}}}{\text{D{\scriptsize ECIDE}}}{\text{D{\scriptscriptstyle ECIDE}}}{\text{DECIDE}}}DECIDE. After the tree absorbs the new evidence, the coordinator decides whether to continue expanding a direction, prune a falsified subtree, stop the run, or attempt to merge a candidate branch. Promotion is guarded by a held-out merge gate: the candidate is evaluated on Etest\mathcal{E}_{\mathrm{test}}Etest in a fresh worktree and is merged into Mbest\mathcal{M}_{\mathrm{best}}Mbest only if it improves over the current best under the objective O\mathcal{O}O. This gate separates exploratory success on Edev\mathcal{E}_{\mathrm{dev}}Edev from verified artifact-level progress.
4.3.2 Executor: Hypothesis-Bound Experimentation
An executor implements one local experiment for one assigned hypothesis. Given a node nnn, it receives the hypothesis hnh_nhn, relevant ancestor insights, the current best artifact, and the development evaluator Edev\mathcal{E}_{\mathrm{dev}}Edev. It then creates an isolated worktree, applies the minimal intervention needed to realize hnh_nhn, runs the evaluator, inspects failures or inactive code paths, and repairs its own implementation when necessary. This local loop may involve multiple edits and reruns, but it remains bound to the assigned hypothesis.
The executor returns exactly the evidence consumed by the coordinator's tree interface: a comparable dev score for selection, a factual result for future ideation, a distilled insight for backpropagation, and a branch reference for held-out verification. This contract is important. If an executor were allowed to change the hypothesis when the metric stalls, the returned score would no longer be evidence about the assigned node, and ancestor-level insights would become difficult to interpret. By keeping executors hypothesis-bound,
Arbor keeps local engineering flexibility while preserving the semantic meaning of tree updates. Algorithm 1 summarizes the full HTR procedure.5. Experiments
5.1 AO Task Suite
To test whether
Arbor can improve real research artifacts, we first construct several AO tasks from actual research tasks. Each task consists of an initial material M0\mathcal{M}_0M0, a natural language objective O\mathcal{O}O, an executable development evaluator Edev\mathcal{E}_{\mathrm{dev}}Edev, a held-out test evaluator Etest\mathcal{E}_{\mathrm{test}}Etest, and a task-native metric. Table 1 gives a compact summary; the task details are described below.Model training. The model-training tasks evaluate whether an agent can improve training algorithms under expensive experimental feedback. In Optimizer Design, we use NanoGPT-Bench ([17]), a benchmark for accelerating NanoGPT training. The initial material is the official tuned Muon optimizer baseline distributed with NanoGPT-Bench, and the objective is to reach the target NanoGPT validation loss in as few optimization steps as possible. The development evaluator uses the standard NanoGPT-Bench task during search, while the test evaluator reruns the selected optimizer with two held-out random seeds and reports the average number of steps. In Architecture Design, we use the
autoresearch benchmark ([18]). The agent modifies a given LLM training codebase, with the goal of obtaining a lower final loss under a fixed time budget. The test evaluator again averages two held-out random-seed runs.Harness engineering. The harness-engineering tasks evaluate whether an agent can improve the control logic around another agent. In Terminal-Bench 2.0, the initial material is the standard official terminal-agent codebase for Terminal-Bench 2.0 ([19]), and the objective is to improve pass rate on terminal-based code and shell tasks. We stratify the 89 tasks by difficulty into 36 development tasks and 53 held-out test tasks, rather than optimizing on the full benchmark. In BrowseComp, the initial material is our standard minimal ReAct-style search harness ([4, 20]). The objective is to improve answer accuracy on browsing questions; the development and test sets are 50 and 300 non-overlapping BrowseComp questions, respectively.
Data synthesis. The data-synthesis tasks evaluate whether an agent can improve a generation pipeline whose output is judged by downstream model behavior. In Search-Agent Data Synthesis, the initial material is a hand-designed pipeline for generating search-agent questions from seed knowledge. Development uses 50 seed items and test uses 100 disjoint seed items. In Math-Reasoning Data Synthesis, the initial material is a hand-designed pipeline for generating AIME-style reasoning problems; development generates 50 problems with 10 seed and test generates 96 problems with 12 seed. Both tasks are scored by the mean pass@4−pass@1\mathrm{pass@4}-\mathrm{pass@1}pass@4−pass@1 gap under a strong GPT-5.5-based ReAct evaluator. This metric rewards problems that are not solved immediately but can be solved with additional attempts.
5.2 Experimental Setup
Benchmarks. Our evaluation uses two complementary types of benchmarks. The first is the AO Task Suite in Section 5.1, which consists of real research tasks with task-specific materials, objectives, development evaluators, and held-out test evaluators. The second is MLE-Bench Lite, a long-horizon machine learning engineering benchmark derived from MLE-bench ([21]), which allows comparison against established benchmark systems under the official task setup and reporting protocol.
Baselines. For the real research tasks, we compare against two strong coding-agent baselines: Codex ([7]) using GPT-5.5 and Claude Code ([8]) using Claude Opus 4.6. Each baseline receives the same initial material, objective, evaluator, and resource budget as
Arbor, and is allowed to inspect files, edit code, run experiments, and iterate until the budget is exhausted. For MLE-Bench Lite, we compare against reported benchmark systems, including AIDE ([22]), ML-Master ([54]) and ML-Master 2.0 ([45]), AIRA-dojo ([55]), InternAgent ([56]), R&D-Agent ([24]), Famou-Agent 2.0 ([57]), MARS ([29]), Leeroo ([58]), AIBuildAI ([59]), LoongFlow ([25]), and AI-Scientist-style systems ([12, 1]). The baseline numbers in Table 3 are adopted from the official MLE-Bench leaderboard and the AI-Scientist paper [1].Metrics. We report native task metrics in the main results, using the direction indicated in Table 1. For cross-task averages and ablations, we also report a normalized held-out improvement over the initial material after orienting all metrics so larger is better. For Δ\DeltaΔ rows, percentage-valued metrics use absolute changes; non-percentage metrics such as steps and loss use the relative improvement below:
where S~\tilde{S}S~ is the native score for higher-is-better metrics and the negated native score for lower-is-better metrics. To measure reliability, we run each stochastic method three times and report Avg@3 with standard deviation unless otherwise specified. For MLE-Bench Lite, we report the official benchmark metrics, including valid-submission rate, above-median rate, any-medal rate, and medal breakdown.
Implementation details. Unless otherwise noted, both the coordinator and executors use Claude Opus 4.6 as the backbone model. All real-research-task runs, including Codex ([7]), Claude Code ([8]), and
Arbor, use a 48-hour wall-clock limit. To keep the two single-agent baselines running over this long horizon without manual intervention, we launch Codex and Claude Code through their official /goal mode, which lets each agent autonomously sustain a long-running task and avoid mid-trajectory interruptions; Arbor is launched through its own coordinator loop. The default Arbor budget is 20 coordinator cycles with maximum tree depth 2. Executor parallelism is bounded by the available evaluator resources, and all wall-clock time, token usage, and evaluator calls are counted when comparing against baselines. The same prompt-level task description is used across methods; Arbor receives no task-specific search strategy beyond the adapter that runs the evaluator and parses scores. For MLE-Bench Lite, every task is optimized on a single NVIDIA A100 GPU under the official benchmark resource budget.5.3 Main Results on Real Research Tasks
Table 2 compares Arbor with Codex and Claude Code on six real research tasks. We focus on two observations.
Arbor gives stronger and more general held-out gains. Arbor obtains the best held-out result on all six tasks, covering three different types of research artifacts: training algorithms, agent harnesses, and data-generation pipelines. The same controller and hypothesis-tree depth are used across these tasks; only the initial material and evaluator are changed. This suggests that the improvement comes from the search procedure itself rather than from task-specific tuning. The gains are also larger than those of single-trajectory coding agents. On BrowseComp, Arbor improves held-out accuracy from 45.3345.3345.33 to 67.6767.6767.67, while Codex and Claude Code reach 50.0050.0050.00 and 53.3353.3353.33. On Math-Reasoning Data Synthesis, Arbor improves the held-out pass-gap by 19.7919.7919.79 points, compared with 5.215.215.21 and 7.297.297.29 points for Codex and Claude Code. The baselines can still make progress, but their gains are smaller and less stable across task types. This supports our main hypothesis: in AO, the bottleneck is not only local code editing, but organizing many trials into a coherent exploration process. The cost results in Section 5.8 further show that Arbor achieves these gains without relying on substantially larger token budgets.
The dev/test split exposes overfitting during autonomous search. Development feedback is useful for guiding exploration, but it is not a reliable admission criterion. Because the agent repeatedly optimizes against EdevE_{\mathrm{dev}}Edev, it can overfit to the development split or exploit evaluator-specific patterns. This is especially visible on Terminal-Bench: Claude Code achieves the highest development score (75.0075.0075.00), but its held-out score drops to 71.7071.7071.70; Arbor has a lower development score (72.2272.2272.22), but reaches the best held-out score (77.3677.3677.36). This gap motivates the held-out merge gate in Arbor. We use EdevE_{\mathrm{dev}}Edev to guide hypothesis search, but promote a candidate artifact only when it improves EtestE_{\mathrm{test}}Etest. This separates exploratory feedback from verified progress. It also makes dev/test disagreement informative: a high-dev, low-test candidate is treated not as a success, but as evidence that the current direction may be exploiting the feedback signal rather than producing a transferable improvement.
5.4 Results on MLE-Bench Lite
We also evaluate
Arbor on MLE-Bench Lite under the official protocol. Unlike our AO task suite, this benchmark fixes the competition-style ML tasks, scoring rules, and medal thresholds, so it tests whether the same controller can turn repeated experiments into stronger runnable submissions. Arbor uses the same controller as before, adding only an adapter for workspace setup and submission formatting. Table 3 reports the results.With a matched Gemini-3-Flash backbone,
Arbor reaches 100% valid submissions, 86.36% above-median rate, and 81.82% any-medal rate, tying the best same-backbone any-medal result while obtaining a higher gold rate than AI-Scientist and LoongFlow. Replacing the backbone with GPT-5.5, without changing the controller, depth, scheduler, or adapter, further raises any-medal to 86.36% and gold to 77.27%, the highest values in Table 3. These results suggest that the hypothesis-tree organization transfers beyond our constructed AO tasks to established long-horizon ML engineering benchmarks.5.5 Backbone Generality
We next test whether
Arbor's gains are tied to a particular backbone model. We repeat representative runs with different backbones. As shown in Figure 3(a), Arbor is not tied to a single frontier model: even with Gemini-3-Flash, a lighter backbone than the Claude and GPT variants, the same controller still improves both browsecomp and MLE-Bench Lite. This suggests that HTR{\mathchoice{\text{H{\scriptsize TR}}}{\text{H{\scriptsize TR}}}{\text{H{\scriptscriptstyle TR}}}{\text{HTR}}}HTR provides a model-agnostic structure for exploration and memory rather than depending on a specific model.We also notice that backbone effects are task-dependent. Although HTR{\mathchoice{\text{H{\scriptsize TR}}}{\text{H{\scriptsize TR}}}{\text{H{\scriptscriptstyle TR}}}{\text{HTR}}}HTR provides a common structure for exploration and memory, final performance is mediated by the compatibility between a model's capabilities and the task requirements. Claude Opus 4.6 performs best on BrowseComp, where improving a search harness relies heavily on broad reasoning and error diagnosis. In contrast, GPT-5.5 performs best on MLE-Bench Lite, where gains are more closely tied to ML-engineering knowledge, including data processing, training recipes, and leaderboard-oriented optimization. Thus,
Arbor is model-agnostic at the framework level, but its empirical ceiling depends on task–backbone fit.5.6 Cross-Task Idea Transfer
A stronger test of generality is whether an optimized artifact transfers beyond the benchmark used for search. This is important for auto-research and auto-harness systems: an agent may improve on a source evaluator by exploiting benchmark-specific patterns rather than discovering generally useful design changes.
We therefore evaluate transfer in the harness-engineering setting.
Arbor is first run on BrowseComp, using only BrowseComp development feedback to propose, implement, and merge search-harness changes. After the run, we freeze the resulting harness and evaluate it directly on two unseen search-agent tasks, HLE and DeepSearchQA, without further task-specific optimization.Figure 3(b) shows that the learned harness transfers. The optimized harness improves BrowseComp held-out accuracy from 45.33% to 67.67%. More importantly, the same frozen codebase also improves HLE from 25.50% to 31.50% and DeepSearchQA from 61.00±6.76%61.00\pm6.76\%61.00±6.76% to 69.00±6.41%69.00\pm6.41\%69.00±6.41%. Since these two tasks are never used during BrowseComp optimization, the gains indicate that Arbor can discover harness-level changes that survive a shift in task distribution, rather than only fitting the source benchmark.
5.7 Ablations
We ablate the two components most central to HTR on MLE-Bench Lite: the hierarchical hypothesis tree and insight feedback. The w/o tree variant reduces search to a flat experiment queue, with all experiments attached directly to the root. The w/o insight feedback variant keeps the tree structure but disables upward propagation of distilled lessons. Both variants use the same tool access, workspace budget, evaluation protocol, and Claude Opus 4.6 backbone as the full system.
HTR improves refinement rather than executability. Table 4 shows that all variants obtain 100% valid submissions, indicating that the ablation gap is not caused by basic execution failure. The difference instead appears in outcome quality. Full Arbor reaches 81.82% Any Medal, compared with 63.64% for w/o tree and 54.54% for w/o insight feedback. The same pattern appears in stronger categories such as Above Median, Silver, and Gold. This suggests that HTR mainly improves later-stage research refinement: once a runnable solution exists, the tree helps the agent decide which directions to extend, revise, or abandon.
The tree is useful only when evidence can accumulate over it. Removing insight feedback while keeping the tree causes a larger drop than removing the tree entirely. This result suggests that hierarchy alone is not sufficient. A tree without propagated lessons can still organize experiments syntactically, but it does not provide the semantic memory needed for later decisions. In contrast, full Arbor uses the tree as a substrate for accumulating evidence: leaf-level results are abstracted into direction-level lessons, which then constrain future ideation and selection.
Tree structure and insight feedback are complementary. The full system outperforms both ablations, indicating that the two components address different parts of the search problem. The tree defines where competing hypotheses are stored and compared, while insight feedback determines what reusable information is carried forward. Their combination allows Arbor to convert local experimental outcomes into persistent constraints on future search, rather than treating each experiment as an isolated trial.
5.8 Token Consumption and Search Cost
We further examine whether
Arbor's gains mainly come from increased model budget. Figure 4 reports total token consumption and relative held-out gain, while Table 5 summarizes the corresponding tree traces.Structured search rather than larger sampling. Across the six completed cost logs,
Arbor uses 20.12M–43.19M tokens, a comparable scale to the single-trajectory baselines. Within this budget, Arbor achieves larger held-out gains on most tasks. This suggests that the improvement is not simply due to spending substantially more tokens, but to how the budget is organized: tokens are used to maintain competing hypotheses, run isolated executions, compare evidence, and update the search tree.Dev improvements are filtered by held-out admission. Table 5 also shows that many nodes improve the development score, but only a smaller subset are merged. This gap is expected. A dev-improving node may still be worse than the current best artifact, or may overfit the development evaluator and fail to transfer to the held-out test. The merge gate therefore prevents local development gains from being mistaken for artifact-level progress. In this sense, the tree trace records broad exploration, while the held-out gate admits only verified improvements into the final artifact.
6. Discussion
We analyze
Arbor's internal research traces to understand how autonomous research progresses once the agent starts running experiments. We focus on three questions: how hypotheses change over time (Section 6.1), when useful improvements appear (Section 6.2), and what kinds of ideas the Hypothesis Tree produces (Section 6.3).6.1 Hypothesis Refinement Analysis
We analyze the BrowseComp hypothesis tree, reporting the main hypothesis shifts, the nodes that triggered them, and the final design selected by the merge gate. Figure 6 traces all three contractions of task understanding alongside the experimental nodes that drove each transition. We find that:
Early nodes test whether a broad mechanism holds. The run begins from a coarse hypothesis and uses the first experiments to confirm or reject it. In BrowseComp, the initial hypothesis is that search agents produce near-miss answers by matching salient cues while missing fine-grained constraints; constraint-decomposed verification and hostile-contradiction checking both improve development accuracy, confirming that fine-grained answer checking is a valid source of gain.
Later nodes localize the bottleneck by probing the mechanism's boundary. Once a mechanism is confirmed, the tree tests where it stops working rather than pushing it further. In BrowseComp, the verifier nodes can judge candidates produced by the search process but rarely recover candidates that were never surfaced, and part of the hostile-verifier gain comes from answer normalization rather than reliable evidence discovery. This shifts the main design target from stricter verification to broader evidence coverage.
Ancestor insights compress these results into the constraints that shape the final design. The accumulated positive and negative findings define what the successful design must satisfy. In BrowseComp, the evidence-dossier aggregator preserves candidates and supporting evidence across independent rollouts, recovering correct answers that appear in only a minority of trajectories; follow-up nodes then rule out persona-diverse rollouts (which only rerank within the same retrieval frontier), the search-augmented judge (which overfits development questions), and shared decomposition (which reduces trajectory independence).
Arbor thus learns that BrowseComp benefits from sharing evidence while keeping search trajectories independent.Takeaway. Hypothesis refinement in HTR is a deepening of task understanding: early nodes test broad mechanisms, later nodes identify their limits, and ancestor insights summarize these results into constraints for the next round of proposals. This constraint accumulation is the core process-level benefit over flat trial-and-error.
6.2 Search Efficiency Analysis
We analyze when the best candidates appear during a run, using each node's execution time and development gain. We find that:
Strong candidates often appear after the search state has accumulated constraints. Across tasks,
Arbor frequently reaches its best candidate in the middle or later part of the run. These improvements are supported by earlier nodes that identify useful mechanisms, rule out weak variants, and narrow the design space.Later proposals are more targeted than early proposals. In BrowseComp, the final evidence-sharing design appears only after several failed or partially successful verifier variants establish that fine-grained checking matters, that candidate coverage is the bottleneck, and that judge-side search and shared decomposition introduce new failure modes. The final proposal is therefore generated from a more informative research state than the initial proposals.
The tree improves search by changing the proposal distribution over time.
Arbor does not simply allocate budget uniformly over independent attempts: its later nodes are conditioned on accumulated evidence from ancestors and siblings. Successful mechanisms become priors, failed variants become negative constraints, and partial gains become starting points for refined hypotheses.Takeaway.
Earlier experiments persistently reduce the arbitrariness of later search, placing mid-to-late improvements on a higher information baseline. The relevant notion of efficiency is whether the same budget produces a less repetitive and more constrained evidence chain, rather than running long enough to stumble onto a result by chance.
6.3 Idea Quality Analysis
We analyze representative ideas generated across model training, harness engineering, and data synthesis tasks, classifying each by its granularity, implementation target, and relation to previous evidence. Figure 7 shows representative examples. We find that:
Most useful ideas are local and executable. In model training, ideas usually modify a specific optimizer component, training recipe, or architecture choice. In harness engineering, they change concrete parts of the agent loop, such as retrieval, aggregation, verification, or context management. In data synthesis, they refine generation, filtering, difficulty calibration, or verification modules. This locality makes each idea easy to implement, evaluate, and attribute to a tree node.
Useful ideas are often evidence-conditioned. Many successful proposals directly respond to earlier observations: the BrowseComp evidence-dossier design follows from the failure mode of verifier-only approaches, and similar patterns appear in data synthesis, where later nodes repair specific weaknesses in difficulty calibration or answer verification. HTR therefore helps convert local failures into new design constraints, and ensures that "half-right" results become the starting point for a more precise hypothesis rather than a reason to abandon the direction.
High-level problem formulation remains important.
Arbor is strongest when the objective can be improved through a sequence of concrete refinements, and less reliable when progress requires a new high-level formulation weakly connected to the existing tree. The Architecture Design task ultimately acknowledged that single-knob tuning had reached diminishing returns and a larger algorithmic move was needed, but identifying that move still depended on prior judgment rather than anything the tree could automatically generate. This highlights the role of human-provided task design: the initial artifact, evaluator, metric, and search interface shape the kinds of ideas the agent can discover.Takeaway. As the tree grows, what has been ruled out, validated, and found to have boundary conditions all become priors constraining the next round of proposals.
Arbor's ideas are therefore not isolated guesses but local advances relative to a known task understanding. Together with the refinement and timing results above, this paints the complete picture of HTR as a process mechanism that makes autonomous research cumulative: not more attempts, but less repetitive and more memory-aware search.7. Conclusion
We presented
Arbor as a framework for Autonomous Optimization, where a research agent must improve a real artifact through long-horizon experimental feedback rather than execute a single predefined trajectory. The core idea is to make the research state persistent and operational: Arbor represents competing hypotheses, artifact versions, evaluation results, failure attributions, and reusable insights in a durable hypothesis tree. A coordinator uses this tree to manage strategic search, while short-lived executors ground individual hypotheses in isolated worktrees and return structured evidence. Together with insight propagation and a held-out admission gate, this design turns trial and error into an auditable process of branching, falsification, and evidence-constrained improvement.Across the AO settings studied here, this organization provides consistent evidence of value. On six real-research tasks spanning model training, harness engineering, and data synthesis,
Arbor achieves the strongest held-out results among the compared methods; on MLE-Bench Lite, the same controller transfers to an established long-horizon ML-engineering benchmark. The transfer study shows that a BrowseComp-optimized harness can improve unseen search-agent tasks, and the ablations indicate that the hypothesis tree and insight feedback are most useful when they operate together. These results support the view that persistent hypothesis management is a useful abstraction for autonomous research, while the limitations of the current task suite, scalar objectives, model capabilities, and search cost leave substantial room for broader and more rigorous future evaluations.Appendix
\setupappendixtoc
A. Limitations and Future Work
Although the empirical results demonstrate the promise of
Arbor for autonomous research, this study has several limitations. We discuss these limitations below and outline the corresponding directions for future work.Evaluation scope. Our experiments are an initial probe of autonomous research rather than a complete benchmark for scientific discovery. The current AO task suite covers model training, harness engineering, and data synthesis, but it does not yet span the full diversity of research problems. Within AI, future tasks should include settings such as low-level kernel optimization, pretraining data-mixture design, and more open-ended system design. Beyond AI, domains such as biology, mathematics, and physics require benchmarks where valuable hypotheses are harder to specify and validate. A broader suite should therefore evaluate not only metric improvement, but also whether the generated ideas are scientifically meaningful, reproducible, and transferable.
Objective design. The present AO interface mainly optimizes a fixed scalar objective defined by a task-specific evaluator. This is useful for controlled experiments, but it is a simplification of real research. Scientific objectives are often multi-dimensional: performance, resource use, robustness, interpretability, novelty, and safety may all matter, and improving one can hurt another. Future AO systems should support multi-objective search, explicit constraints, Pareto-style comparison, and adaptive scheduling between competing criteria. This would also reduce the risk that an agent overfits to a narrow benchmark metric while missing the broader research goal.
Idea generation. We observe that agents can read evaluation feedback carefully and propose useful local refinements, but their research ability remains far from that of expert human researchers. In difficult tasks, they may fail to identify a genuinely new mechanism, abandon a promising direction after early failures, or reverse-engineer solutions from observed scores instead of reasoning from first principles. A more fine-grained study of agent idea formation is therefore needed. Promising directions include better uncertainty tracking, explicit reuse of negative evidence, mechanisms for revisiting suspended branches, and training or prompting methods that encourage causal and first-principles hypotheses rather than only result-driven fixes.
Cost and infrastructure. Long-horizon autonomous research is limited not only by idea quality, but also by systems engineering. In our runs, performance and efficiency depend on details such as prompt caching, evaluator scheduling, isolated environment startup, parallel worktree execution, and the reliability of inter-agent coordination. Large numbers of model calls, evaluator calls, and artifact handoffs can make a successful search expensive even when each individual step is simple. Future work should develop cost-aware tree policies, adaptive evaluator allocation, stronger caching and checkpointing, and more robust execution infrastructure so that AO systems can scale without turning search breadth into uncontrolled compute cost.
Model capability. Finally,
Arbor inherits the strengths and weaknesses of the underlying LLMs. Current models are often capable of coding, summarizing results, and making plausible local hypotheses, but they can still struggle with deep domain knowledge, long chains of causal reasoning, and genuinely creative problem reformulation. Stronger foundation models will likely improve AO directly, but model scaling alone may not be sufficient. Future systems should combine LLMs with domain knowledge bases, specialized tools, simulators, formal checkers, and training signals targeted at scientific hypothesis generation. In this sense, Arbor provides a structure for accumulating and testing ideas, while the quality of those ideas remains an important frontier.B. Details of Arbor
Figure 8 summarizes the implementation-level agent internals used by
Arbor.B.1 Prompts
B.1.1 Coordinator Prompt
B.1.2 Executor Prompt
B.2 Algorithm Workflow
Algorithm 2 expands the HTR pseudocode from the main paper (Algorithm 1) with implementation-level detail. Notation follows the main paper: ιn\iota_nιn denotes the insight recorded at node nnn, bnb_nbn the git branch reference, sns_nsn the Edev\mathcal{E}_\mathrm{dev}Edev score, rnr_nrn the factual result, and ιanc(n)\iota_{\mathrm{anc}(n)}ιanc(n) the concatenated insights on the path from n0n_0n0 to nnn's parent.
lt;$200-word) summary. Leaf insights describe concrete implementations, parent insights summarize families of interventions, and the root insight maintains a global understanding of the problem. This layerwise abstraction lets the coordinator reason at the right granularity without rereading raw logs. The concrete schema and persistence format are detailed in Appendix B.5.
**Experiment management.** Each pending node is executed by an executor dispatched through `RunSubagent` or `RunSubagentParallel`. Rather than editing the shared working tree in place, `Arbor` creates a fresh git worktree branched from the current trunk `HEAD` for each candidate, so every hypothesis gets a clean, independently recoverable experimental boundary and several executors can edit overlapping files concurrently without corrupting the trunk or one another. At launch the executor is injected with the assigned hypothesis, the ancestor insights along its path, and the task objective, development evaluator $\mathcal{E}_\mathrm{dev}$, held-out evaluator $\mathcal{E}_\mathrm{test}$, metric direction, data split, and baseline score stored in the tree metadata via `TreeSetMeta`. Its permission boundary is deliberately narrow: the executor's tools (Table 8) act *only* within its own worktree. It cannot read the trunk, inspect sibling branches, mutate the tree, or run $\mathcal{E}_\mathrm{test}$. It may repair implementation bugs and choose reasonable engineering details, but it may not swap the assigned hypothesis for a different research direction, which keeps the returned evidence attributable to the node actually tested. The executor reports back four separated fields: the development score $s_n$, a factual `result` $r_n$ (raw observations such as errors, curves, and metric breakdowns), a distilled `insight` $\iota_n$ (the causal lesson), and the branch reference $b_n$. The coordinator auto-extracts these, updates the node, and decides admission: a result counts as a *useful gain* only when it improves $\mathcal{E}_\mathrm{dev}$ over the trunk by at least the merge threshold, which triggers the `GitMergeBranch` held-out gate that re-runs $\mathcal{E}_\mathrm{test}$ in a separate detached worktree and promotes the artifact only if it strictly beats the current best. Every other outcome is treated as *failure evidence* rather than noise: the node's insight and, if the direction is abandoned, its `TreePrune` reason are aggregated into the constraint view (`TreeView(format="constraints")`) that conditions the next $\textsc{Ideate}$ step, so negative results actively narrow the search space. Tool-level details of dispatch, score extraction, template-variable substitution (`{cwd}`, `node_id`), and artifact filtering are given in Appendix B.4.
**Long-horizon operation.** A single auto-research run can span hundreds of turns and many hours of evaluation, far beyond one context window, so `Arbor` adds explicit mechanisms to keep the search productive instead of stalling on early failures or chasing noisy evaluation swings. First, because all durable state lives in the persisted idea tree rather than the transcript, the run survives crashes, agent restarts, and context compression; post-commit context pruning further elides spent $\textsc{Ideate}$ scratch work and loaded skill bodies once a candidate is committed, bounding context growth. Second, long training or evaluation commands are routed through `RunTraining`, which blocks until completion or timeout while continuously capturing partial metrics, progress logs, and checkpoints, so even a run that times out at 80% of its epochs returns actionable evidence instead of a silent failure. Third, a *convergence detector* monitors recent score velocity over a sliding window and counts consecutive non-improving experiments, escalating through `warn`, `paradigm_shift`, and `stop` signals; it also flags *parent exhaustion* when a parent's recent children all fail to beat the trunk, prompting the coordinator to summarize the failure pattern and open a fresh depth-1 direction rather than over-exploring a dead branch. A meaningful-improvement threshold prevents a single noisy uptick from resetting this signal, balancing premature stagnation against unbounded exploration.
**Functional extensibility (skills and plugins).** To remain adaptable across research domains without changing the core loop, `Arbor` exposes two extension surfaces. *Skills* are markdown documents with YAML frontmatter, discovered by a registry from both the built-in directory and a project-local `.research_agent/skills/` override, and loaded on demand via `LoadSkill`. The $\textsc{Ideate}$ protocol, for instance, loads `idea_drafting`, `first_principles_probe`, and `fatal_flaw_scan` before proposing candidates, then prunes their bodies from context after each commit, so reasoning guidance is injected just-in-time and does not permanently inflate context. *Plugins* are YAML domain adapters that specialize the system declaratively rather than through code: each plugin can inject domain guidance at six prompt points (coordinator/executor init, ideate, decide, preamble, and workflow), declare an evaluation contract, mark protected paths and required outputs for the merge guard, override runtime configuration through named profiles, and tune convergence thresholds. For example, the `mle_kaggle` plugin configures a Kaggle/MLE-bench task with its metric direction, evaluation command, protected data paths, and time-budget profiles entirely in YAML. Together, skills and plugins let `Arbor` adapt to new artifacts and evaluation regimes while keeping the hypothesis-tree machinery and agent contracts unchanged.
#### B.4 Agent Tools and Hyperparameter Settings
##### B.4.1 Coordinator Tools
The coordinator's tool set reflects a strict division of labor: it may read any file in the repository and inspect the tree in any projection, but it may never edit code directly. All implementation work is delegated to executors via `RunSubagent` or `RunSubagentParallel`. This design is intentional. If the coordinator could edit code directly, it would be tempted to make small local fixes without creating a new tree node, which would break the invariant that every code change is traceable to a hypothesis. Enforcing the edit boundary at the tool level makes this invariant a system property rather than a behavioral expectation.
Table 6 lists the full tool set. The tree tools (`TreeView`, `TreeAddNode`, `TreeUpdateNode`, `TreePrune`, `TreeSetMeta`, `TreePropagate`) are the primary interface through which the coordinator manages the shared research state. The dispatch tools (`RunSubagent`, `RunSubagentParallel`) are the boundary through which the coordinator hands off implementation to executors and receives structured evidence back. The merge tool (`GitMergeBranch`) is the only path through which a candidate branch can be promoted to the current best, enforcing the held-out gate at the tool level rather than relying on the coordinator to remember to verify before merging. Table 7 reports the key hyperparameters used across all experiments.
" data-original-markdown="
#### B.3 Key Design of `Arbor` Framework
General-purpose coding agents such as Codex and Claude Code are built for general-purpose software tasks: they chain tool calls on a single working tree to edit, test, and fix code against a goal that is already well specified. `Arbor` instead targets the auto-research setting, and we make a series of engineering choices that specialize the agent for it so that the system can flexibly adapt to different research needs and manage many experiments over a long horizon. We group its key engineering designs into four areas: (i) hypothesis-tree management, (ii) experiment management, (iii) long-horizon operation, and (iv) functional extensibility through skills and plugins.
**Hypothesis-tree management.** Unlike a code agent whose memory is the linear chat transcript, `Arbor` externalizes its research state into an explicit hypothesis tree (the *idea tree*), an in-memory object that is the single authoritative record of a run. Each node holds a hierarchical dotted address (e.g. `ROOT`, `1`, `1.1`) that encodes its path from the root, its `parent_id` and `children_ids`, a once-written `hypothesis`, a `status` field tracing the lifecycle `pending` $\to$ `running` $\to$ `done` $\to$ `merged`, `pruned`, a development `score`, a factual `result`, a distilled `insight`, and a `code_ref` branch pointer to the artifact. Depth-1 nodes are broad directions and deeper nodes are concrete refinements, so edges encode hypothesis refinement rather than chronological actions. The coordinator never touches the raw structure; it operates the tree only through a small set of typed tools (`TreeAddNode`, `TreeUpdateNode`, `TreePrune`, `TreeSetMeta`, `TreePropagate`) and read projections (`TreeView`). Storing an experiment is thus a controlled mutation: the executor's structured report is parsed and written into the node's `score`/`result`/`insight`/`code_ref` fields, and the tree is serialized to JSON (and rendered to Markdown) after *every* mutation. Crucially, when a node finishes `Arbor` *backpropagates* its insight: `propagate_insights` walks from the node's parent to the root and, at each ancestor, an LLM synthesizes the insights of that ancestor's children into a concise (lt;$200-word) summary. Leaf insights describe concrete implementations, parent insights summarize families of interventions, and the root insight maintains a global understanding of the problem. This layerwise abstraction lets the coordinator reason at the right granularity without rereading raw logs. The concrete schema and persistence format are detailed in Appendix B.5.
**Experiment management.** Each pending node is executed by an executor dispatched through `RunSubagent` or `RunSubagentParallel`. Rather than editing the shared working tree in place, `Arbor` creates a fresh git worktree branched from the current trunk `HEAD` for each candidate, so every hypothesis gets a clean, independently recoverable experimental boundary and several executors can edit overlapping files concurrently without corrupting the trunk or one another. At launch the executor is injected with the assigned hypothesis, the ancestor insights along its path, and the task objective, development evaluator $\mathcal{E}_\mathrm{dev}$, held-out evaluator $\mathcal{E}_\mathrm{test}$, metric direction, data split, and baseline score stored in the tree metadata via `TreeSetMeta`. Its permission boundary is deliberately narrow: the executor's tools (Table 8) act *only* within its own worktree. It cannot read the trunk, inspect sibling branches, mutate the tree, or run $\mathcal{E}_\mathrm{test}$. It may repair implementation bugs and choose reasonable engineering details, but it may not swap the assigned hypothesis for a different research direction, which keeps the returned evidence attributable to the node actually tested. The executor reports back four separated fields: the development score $s_n$, a factual `result` $r_n$ (raw observations such as errors, curves, and metric breakdowns), a distilled `insight` $\iota_n$ (the causal lesson), and the branch reference $b_n$. The coordinator auto-extracts these, updates the node, and decides admission: a result counts as a *useful gain* only when it improves $\mathcal{E}_\mathrm{dev}$ over the trunk by at least the merge threshold, which triggers the `GitMergeBranch` held-out gate that re-runs $\mathcal{E}_\mathrm{test}$ in a separate detached worktree and promotes the artifact only if it strictly beats the current best. Every other outcome is treated as *failure evidence* rather than noise: the node's insight and, if the direction is abandoned, its `TreePrune` reason are aggregated into the constraint view (`TreeView(format="constraints")`) that conditions the next $\textsc{Ideate}$ step, so negative results actively narrow the search space. Tool-level details of dispatch, score extraction, template-variable substitution (`{cwd}`, `node_id`), and artifact filtering are given in Appendix B.4.
**Long-horizon operation.** A single auto-research run can span hundreds of turns and many hours of evaluation, far beyond one context window, so `Arbor` adds explicit mechanisms to keep the search productive instead of stalling on early failures or chasing noisy evaluation swings. First, because all durable state lives in the persisted idea tree rather than the transcript, the run survives crashes, agent restarts, and context compression; post-commit context pruning further elides spent $\textsc{Ideate}$ scratch work and loaded skill bodies once a candidate is committed, bounding context growth. Second, long training or evaluation commands are routed through `RunTraining`, which blocks until completion or timeout while continuously capturing partial metrics, progress logs, and checkpoints, so even a run that times out at 80% of its epochs returns actionable evidence instead of a silent failure. Third, a *convergence detector* monitors recent score velocity over a sliding window and counts consecutive non-improving experiments, escalating through `warn`, `paradigm_shift`, and `stop` signals; it also flags *parent exhaustion* when a parent's recent children all fail to beat the trunk, prompting the coordinator to summarize the failure pattern and open a fresh depth-1 direction rather than over-exploring a dead branch. A meaningful-improvement threshold prevents a single noisy uptick from resetting this signal, balancing premature stagnation against unbounded exploration.
**Functional extensibility (skills and plugins).** To remain adaptable across research domains without changing the core loop, `Arbor` exposes two extension surfaces. *Skills* are markdown documents with YAML frontmatter, discovered by a registry from both the built-in directory and a project-local `.research_agent/skills/` override, and loaded on demand via `LoadSkill`. The $\textsc{Ideate}$ protocol, for instance, loads `idea_drafting`, `first_principles_probe`, and `fatal_flaw_scan` before proposing candidates, then prunes their bodies from context after each commit, so reasoning guidance is injected just-in-time and does not permanently inflate context. *Plugins* are YAML domain adapters that specialize the system declaratively rather than through code: each plugin can inject domain guidance at six prompt points (coordinator/executor init, ideate, decide, preamble, and workflow), declare an evaluation contract, mark protected paths and required outputs for the merge guard, override runtime configuration through named profiles, and tune convergence thresholds. For example, the `mle_kaggle` plugin configures a Kaggle/MLE-bench task with its metric direction, evaluation command, protected data paths, and time-budget profiles entirely in YAML. Together, skills and plugins let `Arbor` adapt to new artifacts and evaluation regimes while keeping the hypothesis-tree machinery and agent contracts unchanged.
#### B.4 Agent Tools and Hyperparameter Settings
##### B.4.1 Coordinator Tools
The coordinator's tool set reflects a strict division of labor: it may read any file in the repository and inspect the tree in any projection, but it may never edit code directly. All implementation work is delegated to executors via `RunSubagent` or `RunSubagentParallel`. This design is intentional. If the coordinator could edit code directly, it would be tempted to make small local fixes without creating a new tree node, which would break the invariant that every code change is traceable to a hypothesis. Enforcing the edit boundary at the tool level makes this invariant a system property rather than a behavioral expectation.
Table 6 lists the full tool set. The tree tools (`TreeView`, `TreeAddNode`, `TreeUpdateNode`, `TreePrune`, `TreeSetMeta`, `TreePropagate`) are the primary interface through which the coordinator manages the shared research state. The dispatch tools (`RunSubagent`, `RunSubagentParallel`) are the boundary through which the coordinator hands off implementation to executors and receives structured evidence back. The merge tool (`GitMergeBranch`) is the only path through which a candidate branch can be promoted to the current best, enforcing the held-out gate at the tool level rather than relying on the coordinator to remember to verify before merging. Table 7 reports the key hyperparameters used across all experiments.
" data-source-offset="72466" class="markdown-segment">
B.3 Key Design of Arbor Framework
General-purpose coding agents such as Codex and Claude Code are built for general-purpose software tasks: they chain tool calls on a single working tree to edit, test, and fix code against a goal that is already well specified.
Arbor instead targets the auto-research setting, and we make a series of engineering choices that specialize the agent for it so that the system can flexibly adapt to different research needs and manage many experiments over a long horizon. We group its key engineering designs into four areas: (i) hypothesis-tree management, (ii) experiment management, (iii) long-horizon operation, and (iv) functional extensibility through skills and plugins.Hypothesis-tree management. Unlike a code agent whose memory is the linear chat transcript,
Arbor externalizes its research state into an explicit hypothesis tree (the idea tree), an in-memory object that is the single authoritative record of a run. Each node holds a hierarchical dotted address (e.g. ROOT, 1, 1.1) that encodes its path from the root, its parent_id and children_ids, a once-written hypothesis, a status field tracing the lifecycle pending →\to→ running →\to→ done →\to→ merged, pruned, a development score, a factual result, a distilled insight, and a code_ref branch pointer to the artifact. Depth-1 nodes are broad directions and deeper nodes are concrete refinements, so edges encode hypothesis refinement rather than chronological actions. The coordinator never touches the raw structure; it operates the tree only through a small set of typed tools (TreeAddNode, TreeUpdateNode, TreePrune, TreeSetMeta, TreePropagate) and read projections (TreeView). Storing an experiment is thus a controlled mutation: the executor's structured report is parsed and written into the node's score/result/insight/code_ref fields, and the tree is serialized to JSON (and rendered to Markdown) after every mutation. Crucially, when a node finishes Arbor backpropagates its insight: propagate_insights walks from the node's parent to the root and, at each ancestor, an LLM synthesizes the insights of that ancestor's children into a concise (<<<200-word) summary. Leaf insights describe concrete implementations, parent insights summarize families of interventions, and the root insight maintains a global understanding of the problem. This layerwise abstraction lets the coordinator reason at the right granularity without rereading raw logs. The concrete schema and persistence format are detailed in Appendix B.5.Experiment management. Each pending node is executed by an executor dispatched through
RunSubagent or RunSubagentParallel. Rather than editing the shared working tree in place, Arbor creates a fresh git worktree branched from the current trunk HEAD for each candidate, so every hypothesis gets a clean, independently recoverable experimental boundary and several executors can edit overlapping files concurrently without corrupting the trunk or one another. At launch the executor is injected with the assigned hypothesis, the ancestor insights along its path, and the task objective, development evaluator Edev\mathcal{E}_\mathrm{dev}Edev, held-out evaluator Etest\mathcal{E}_\mathrm{test}Etest, metric direction, data split, and baseline score stored in the tree metadata via TreeSetMeta. Its permission boundary is deliberately narrow: the executor's tools (Table 8) act only within its own worktree. It cannot read the trunk, inspect sibling branches, mutate the tree, or run Etest\mathcal{E}_\mathrm{test}Etest. It may repair implementation bugs and choose reasonable engineering details, but it may not swap the assigned hypothesis for a different research direction, which keeps the returned evidence attributable to the node actually tested. The executor reports back four separated fields: the development score sns_nsn, a factual result rnr_nrn (raw observations such as errors, curves, and metric breakdowns), a distilled insight ιn\iota_nιn (the causal lesson), and the branch reference bnb_nbn. The coordinator auto-extracts these, updates the node, and decides admission: a result counts as a useful gain only when it improves Edev\mathcal{E}_\mathrm{dev}Edev over the trunk by at least the merge threshold, which triggers the GitMergeBranch held-out gate that re-runs Etest\mathcal{E}_\mathrm{test}Etest in a separate detached worktree and promotes the artifact only if it strictly beats the current best. Every other outcome is treated as failure evidence rather than noise: the node's insight and, if the direction is abandoned, its TreePrune reason are aggregated into the constraint view (TreeView(format="constraints")) that conditions the next IDEATE{\mathchoice{\text{I{\scriptsize DEATE}}}{\text{I{\scriptsize DEATE}}}{\text{I{\scriptscriptstyle DEATE}}}{\text{IDEATE}}}IDEATE step, so negative results actively narrow the search space. Tool-level details of dispatch, score extraction, template-variable substitution ({cwd}, node_id), and artifact filtering are given in Appendix B.4.Long-horizon operation. A single auto-research run can span hundreds of turns and many hours of evaluation, far beyond one context window, so
Arbor adds explicit mechanisms to keep the search productive instead of stalling on early failures or chasing noisy evaluation swings. First, because all durable state lives in the persisted idea tree rather than the transcript, the run survives crashes, agent restarts, and context compression; post-commit context pruning further elides spent IDEATE{\mathchoice{\text{I{\scriptsize DEATE}}}{\text{I{\scriptsize DEATE}}}{\text{I{\scriptscriptstyle DEATE}}}{\text{IDEATE}}}IDEATE scratch work and loaded skill bodies once a candidate is committed, bounding context growth. Second, long training or evaluation commands are routed through RunTraining, which blocks until completion or timeout while continuously capturing partial metrics, progress logs, and checkpoints, so even a run that times out at 80% of its epochs returns actionable evidence instead of a silent failure. Third, a convergence detector monitors recent score velocity over a sliding window and counts consecutive non-improving experiments, escalating through warn, paradigm_shift, and stop signals; it also flags parent exhaustion when a parent's recent children all fail to beat the trunk, prompting the coordinator to summarize the failure pattern and open a fresh depth-1 direction rather than over-exploring a dead branch. A meaningful-improvement threshold prevents a single noisy uptick from resetting this signal, balancing premature stagnation against unbounded exploration.Functional extensibility (skills and plugins). To remain adaptable across research domains without changing the core loop,
Arbor exposes two extension surfaces. Skills are markdown documents with YAML frontmatter, discovered by a registry from both the built-in directory and a project-local .research_agent/skills/ override, and loaded on demand via LoadSkill. The IDEATE{\mathchoice{\text{I{\scriptsize DEATE}}}{\text{I{\scriptsize DEATE}}}{\text{I{\scriptscriptstyle DEATE}}}{\text{IDEATE}}}IDEATE protocol, for instance, loads idea_drafting, first_principles_probe, and fatal_flaw_scan before proposing candidates, then prunes their bodies from context after each commit, so reasoning guidance is injected just-in-time and does not permanently inflate context. Plugins are YAML domain adapters that specialize the system declaratively rather than through code: each plugin can inject domain guidance at six prompt points (coordinator/executor init, ideate, decide, preamble, and workflow), declare an evaluation contract, mark protected paths and required outputs for the merge guard, override runtime configuration through named profiles, and tune convergence thresholds. For example, the mle_kaggle plugin configures a Kaggle/MLE-bench task with its metric direction, evaluation command, protected data paths, and time-budget profiles entirely in YAML. Together, skills and plugins let Arbor adapt to new artifacts and evaluation regimes while keeping the hypothesis-tree machinery and agent contracts unchanged.B.4 Agent Tools and Hyperparameter Settings
B.4.1 Coordinator Tools
The coordinator's tool set reflects a strict division of labor: it may read any file in the repository and inspect the tree in any projection, but it may never edit code directly. All implementation work is delegated to executors via
RunSubagent or RunSubagentParallel. This design is intentional. If the coordinator could edit code directly, it would be tempted to make small local fixes without creating a new tree node, which would break the invariant that every code change is traceable to a hypothesis. Enforcing the edit boundary at the tool level makes this invariant a system property rather than a behavioral expectation.Table 6 lists the full tool set. The tree tools (
TreeView, TreeAddNode, TreeUpdateNode, TreePrune, TreeSetMeta, TreePropagate) are the primary interface through which the coordinator manages the shared research state. The dispatch tools (RunSubagent, RunSubagentParallel) are the boundary through which the coordinator hands off implementation to executors and receives structured evidence back. The merge tool (GitMergeBranch) is the only path through which a candidate branch can be promoted to the current best, enforcing the held-out gate at the tool level rather than relying on the coordinator to remember to verify before merging. Table 7 reports the key hyperparameters used across all experiments.












