PATH-GRPO: Placement-Aware Teacher Hinting for Data-Scarce Multi-Turn Agents
Abstract
Training a language-model agent with reinforcement learning (RL) is hard when training tasks are few and most attempts fail. With so few tasks, a supervised warm start before RL has little to learn from. Group-relative methods such as GRPO learn from reward differences within a group of attempts, so a group in which every attempt fails gives no learning signal. Most existing hint methods for such groups study single-turn reasoning, where the hint can only go in the prompt. In a multi-turn agent, a hint can enter at any turn, so where it enters becomes a design choice. We introduce PATH-GRPO (Placement-Aware Teacher Hinting with GRPO). When every attempt fails, a teacher chooses a restart turn, reads the reference solution, and writes an insight that explains how to find the answer. The policy retries from that turn, and we remove the insight before the update (stripping), so that the update sees the same inputs as evaluation, where the policy gets no help. With only 90 AppWorld tasks and 45 TravelPlanner queries, PATH-GRPO improves over GRPO at both 4B and 9B, on AppWorld Test-Challenge scenario goal completion (+13.2 and +9.2 points) and TravelPlanner leaderboard test success (+29.1 and +4.5). Our extensive ablations show that placement matters: on the main metrics, placing the insight at a teacher-chosen restart turn scores higher than placing it in the prompt or at a random turn. For these turn-level retries, both stripping the insight and letting the insight-generator teacher read the reference solution also yield higher scores on these metrics.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.