acceptodds
Under review as a conference paper at ICLR 2027

IterLab: Training Agents to Experiment and Improve over Long Horizons with Calibrated Synthetic Tasks

Abstract

Training agents to improve solutions over long horizons requires tasks that make experimentation both necessary and productive. We introduce IterLab, a pipeline combining calibrated synthetic tasks with supervision on complete experimental histories. Prediction problems with known mechanisms expose structural gains; paired references set quality targets, and recovery checks assess whether input transformations preserve useful signal. Agents inspect data, run code, interpret feedback, and revise solutions within a shared budget. Across 300 tasks, structural features improve reference performance on 214. Quality and interaction screening retain 413 demonstrations from 6,000 requested trajectories. Supervised fine-tuning on these and supplemental demonstrations improves both Qwen3.6-27B and Qwen3.6-35B-A3B on DSBench, MLAgentBench, MLE-bench, and SWE-fficiency. MLE-bench medal rates increase by 15.56 and 14.67 percentage points, respectively, while software-optimization performance improves alongside correctness. Task analyses distinguish modeling from input-processing demands and identify which challenges yield demonstrations. IterLab turns measurable improvement opportunities into training experience for extended, feedback-driven experimentation. Code is available at https://anonymous.4open.science/r/anonymous_code2-CCB1.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.