acceptodds
Under review as a conference paper at ICLR 2027

PATH-GRPO: Placement-Aware Teacher Hinting for Data-Scarce Multi-Turn Agents

Abstract

Training a language-model agent with reinforcement learning (RL) is hard when training tasks are few and most attempts fail. With so few tasks, a supervised warm start before RL has little to learn from. Group-relative methods such as GRPO learn from reward differences within a group of attempts, so a group in which every attempt fails gives no learning signal. Most existing hint methods for such groups study single-turn reasoning, where the hint can only go in the prompt. In a multi-turn agent, a hint can enter at any turn, so where it enters becomes a design choice. We introduce PATH-GRPO (Placement-Aware Teacher Hinting with GRPO). When every attempt fails, a teacher chooses a restart turn, reads the reference solution, and writes an insight that explains how to find the answer. The policy retries from that turn, and we remove the insight before the update (stripping), so that the update sees the same inputs as evaluation, where the policy gets no help. With only 90 AppWorld tasks and 45 TravelPlanner queries, PATH-GRPO improves over GRPO at both 4B and 9B, on AppWorld Test-Challenge scenario goal completion (+13.2 and +9.2 points) and TravelPlanner leaderboard test success (+29.1 and +4.5). Our extensive ablations show that placement matters: on the main metrics, placing the insight at a teacher-chosen restart turn scores higher than placing it in the prompt or at a random turn. For these turn-level retries, both stripping the insight and letting the insight-generator teacher read the reference solution also yield higher scores on these metrics.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.