acceptodds
Under review as a conference paper at ICLR 2027

ScienceFlow: Executable-State Re-Anchoring for Long-Horizon Autonomous Research

Abstract

Long-horizon LLM agents can improve research outcomes, but their gains often rely on expensive model calls and poorly targeted computation. We introduce ScienceFlow, an end-to-end auto-research framework designed to convert limited inference and execution budgets into sustained research progress. ScienceFlow organizes research around recoverable executable states, allowing validated artifacts and intermediate results to persist across iterations. Its Executable-State Transition with Re-Anchoring (ESTRA) mechanism jointly selects an execution anchor from the current or archived states and an extend-or-redirect research direction, enabling continued exploration or redirection from recoverable executable states. An evidence-aware execution controller further admits and monitors jobs according to remaining budget, resource availability, and validated progress. We evaluate ScienceFlow across machine learning, scientific modeling, and mathematical optimization. On the full 75-task MLE-bench, ScienceFlow with DeepSeek-V4-Flash-Preview achieves a Any-Medal rate under a 24-hour budget per task run, evaluated over three independent benchmark sweeps. ScienceFlow also achieves a low mean API expenditure of $76.4 per benchmark sweep under our cost-accounting protocol. The source code is temporarily available in the supplementary material.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.