acceptodds
Under review as a conference paper at ICLR 2027

AI Designs and Trains Tool-Using Agent Teams

Abstract

Developing a tool-using agent requires deciding what to test, how to organize work, and how to learn from execution. These decisions interact: a change in system design alters both the failures that remain and the experience available for training. We introduce ResidualForge, an AI-directed framework that connects executable task construction, failure-guided system design, and native policy training. The designer refines a single actor, tests changes to its skills and division of work, and records which cases each change repairs or regresses. The selected design is then fixed while its policies learn from trajectories that retain actual tool observations and role-to-role handoffs. Across ClawEval, QwenClawBench, and TeamBench, we evaluate complete systems, their development components, and transfer across execution harnesses. On QwenClawBench, on-policy distillation improves the selected Qwen3.5-9B system from 34.36 to 37.42. Removing team handoffs from training context causes the largest component-study drop on ClawEval and QwenClawBench, and pooled training rollouts improve performance in an unseen harness. System selection also retains a single actor on TeamBench when larger candidates perform worse. These results connect autonomous engineering decisions to the quality of the resulting system and the execution experience used to train it.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.