acceptodds
Under review as a conference paper at ICLR 2027

Athanor: An Agent-Native System for Autonomous Post-Training Research

Abstract

Post-training research is an iterative process of diagnosing model weaknesses, testing data and algorithmic recipes, evaluating outcomes, and refining subsequent experiments from evidence. We build , an open agent-native system that enables frontier generalist agents to hill-climb a raw base model on target benchmarks through long-running campaigns of training experiments on real GPUs. The agent develops hypotheses and submits recipes, while system services run training, protect evaluation, and link recipes, code, artifacts, and research notes into a persistent research record. This lets the agent work on the research process itself, from hypothesis formation and experimental design to result analysis and iterative refinement. Across six successive campaigns on eight-GPU nodes, Opus 5 improves Qwen3-4B-Base from to across seven benchmarks using a fixed, decontaminated data pool, against 70.4 for the official instruction-tuned model. We analyze the record of these campaigns: 839 accepted experiments, 711 evaluations, and nearly a thousand research notes. We observe that the agent keeps to a narrow band of recipes, changing one setting at a time over a few trusted corpora. Its largest losses rarely come from failed jobs, but from data it did not inspect and the intrinsic difficulty of post-training: the same recipe can help one checkpoint but hurt another. On the strategy side, it drops early conclusions that later evidence contradicts, while directions it rules out mostly stay ruled out.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.