acceptodds
Under review as a conference paper at ICLR 2027

AutoDataBench: Can Agents Write the Data That Feeds the Self-Improvement Loop?

Abstract

Recent gains in language model capability have come more from data than from architecture. Frontier labs and data companies produce verifiable agentic tasks, which supervised finetuning and reinforcement learning then turn into capability. This production line still rests on human labour and on human-in-the-loop collaboration. Agentic tasks are also the hardest kind of data to produce, because each one needs an executable environment, a reliable verifier, and a difficulty matched to the model being trained. Automating task creation would let data production scale with compute rather than with expert headcount, would extend to more domains, and would enable a key step in recursive self-improvement (RSI). Current evaluations of an agent's ability to write such tasks measure how a model performs after training on what the agent produced. That does not match common practice in the data industry, where a delivery is accepted or rejected one sample at a time, before any training. No existing evaluation asks whether an individual task meets the acceptance criteria of a data pipeline. We therefore introduce AutoDataBench. Given an original benchmark task and a record of the target model attempting it, an agent must write a new task that follows the same benchmark conventions and meets practical acceptance standards on validity, novelty, difficulty and behavioural coverage. We make the job as easy as we can. The agent is given the conventions, a starting task and the behaviour to target, and it needs only to write a task of similar structure that exercises the same behavioural modes. Across three benchmarks of executable agent tasks, no agent we evaluate scores above 20 out of 100 at the default time budget of 45 minutes. Giving the strongest agent four times as long improves its score substantially, while the cost of one usable task stays almost unchanged. Current agents can write training tasks of the required quality, but not efficiently. \sys provides a direct measure of an agent's capacity for autonomous data synthesis: one artifact at a time, judged against the criteria a production pipeline would apply, and without a training run.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.