acceptodds
Under review as a conference paper at ICLR 2027

AutoScaffold: AI-Guided Scaffolding for Reinforcement Learning of Language Models

Abstract

AI-for-AI can improve reinforcement learning (RL) through automated training configuration. We extend this paradigm to adaptive training-time scaffolding with AutoScaffold. A frozen teacher LLM analyzes a student's evolving bottlenecks and provides strategies and hints to generate informative experience when unscaffolded rollouts yield little learning signal. With minimal domain-specific information, the teacher examines training evidence, proposes class-level strategies and task-specific hints, and adjusts their use, guided by held-out validation of class-level revisions. This process requires no hand-crafted curriculum or human intervention during training. Across embodied tasks, search-based question answering, mathematical reasoning, and GPU kernel generation, AutoScaffold improves over corresponding RL baselines when evaluated without scaffolds. Gains include 7.2 percentage points on ALFWorld and 3.8 points in average question-answering exact match. On mathematical reasoning, it surpasses QuestA's human-designed scaffolding strategy with approximately one-third as many student training steps. The teacher also rediscovers intuitive strategies, including withdrawing support as the student learns. These results suggest that adaptive training assistance can convert teacher guidance into the student's own capabilities.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.