Fathom-Autoresearch: Towards Reliable Autonomous Post-Training Agents
Abstract
Autonomous post-training has a fixed deadline but delayed task-level feedback. An experiment can guide the remaining campaign only if its intervention is realized, exposed adequately, and evaluated in time to change the next decision. We pro- pose Fathom-Autoresearch, which couples evidence-grounded research reviews to continuous implementation and checkpoint retention. Across seven PostTrainBench tasks, its selected-checkpoint aggregate is 44.4% for Qwen3-1.7B-Base and 54.4% for Qwen3-4B-Base. It exceeds the persistent Codex CLI-Kimi K3 control in all eight displayed matched task model pairs. Across four tasks per model, effective GPU hours spent on selected-checkpoint work or completed evaluation are 16.7 versus 12.8 at 1.7B and 12.0 versus 8.7 at 4B. A frozen-state replay favors the Fathom-Autoresearch Planner’s written experiments by 2.12 points on a 24-point rubric across 15 states, without testing online execution. Together these layers characterize finite-horizon experimental control rather than checkpoint score mediation by any single role.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.