acceptodds
Under review as a conference paper at ICLR 2027

Toward Generalist Autonomous Research via Hypothesis-Tree Refinement

Abstract

Autonomous research requires agents to make sustained progress over long horizons by continuously refining hypotheses through experimental feedback. Although recent large language model agents can edit code, invoke tools, and execute experiments for extended periods, longer execution alone does not make research cumulative. Without an explicit research state, experiments remain isolated attempts, preventing evidence from being accumulated and reused to guide future exploration. We formulate this setting as Autonomous Optimization (AO), in which an agent iteratively improves an artifact under a fixed objective without step-level human supervision, and introduce Arbor, a general framework for AO. At the core of Arbor is Hypothesis Tree Refinement (HTR), an iterative algorithm that maintains a persistent hypothesis tree as the research state and continuously refines it as isolated experiments return evidence. By grounding each hypothesis in its artifact and evidence, and propagating findings to guide subsequent experiments, HTR connects local outcomes to cumulative hypothesis refinement. The tree serves as the search frontier, long-term memory, and audit trail. Across six real research tasks spanning model training, harness engineering, and data synthesis, Arbor improves the admission metric on every task under the same time budget, yielding roughly twice the gain of Codex and Claude Code. On MLE-Bench Lite it reaches a 77.3% gold rate, 24 points above the strongest prior system in our comparison. After optimization on BrowseComp, the frozen search harness transfers to unseen HLE and DeepSearchQA tasks, improving both without further optimization. Ab- lations and research traces support the value of connecting hypothesis management with evidence propagation for cumulative autonomous research.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.