acceptodds
Under review as a conference paper at ICLR 2027

VibeTraining-Bench: Data–Trainer–Harness Co-Design for Iterative System Improvement

Abstract

Recursive self-improvement requires research agents that can iteratively improve a designated target agent system as a whole, rather than one component in isolation. A target agent system's performance depends jointly on its training data, trainer, and inference harness, and because these components interact, the most useful intervention can shift as the system evolves. Existing benchmarks either fix a single component in advance or open only training-side components. They therefore do not evaluate whether research agents can select, coordinate, and revise interventions across training and inference, a capability we call cross-component co-design. We introduce VibeTraining-Bench,, a benchmark of five tasks spanning terminal interaction, software engineering, multi-turn tool use, and mathematical reasoning. Within each task, research agents start from the same initial target agent system and training setup, then use validation scores and execution trajectories to guide iterative changes to the permitted components under a fixed budget. To limit reward hacking, an isolated grader evaluates the selected target agent system once on hidden test tasks that the research agent cannot access. Evaluating six research-agent configurations, we find large hidden-test gains (32.9% to 53.7% on SWE-bench Verified; 20.3% to 52.1% on -Bench), yet improvement is not guaranteed and component exploration remains uneven: one unrestricted SWE run does not edit the harness, although a separate harness-only run improves the initial target agent system. Validation search is also non-monotonic, and validation leaders do not always lead on the hidden test. These results expose a gap between finding effective system changes and reliably exploring and selecting them.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.