acceptodds
Under review as a conference paper at ICLR 2027

VH-GPO: VERIFIED HANDOFF FOR LLM AGENT REINFORCEMENT LEARNING

Abstract

Reinforcement learning (RL) plays a key role in training large language models (LLMs) to solve complex tasks through outcome-based feedback. Group-relative RL methods such as GRPO train language-model agents by comparing the outcomes of multiple rollouts for the same task. This mechanism is effective when outcomes vary, but if every trajectory in a group fails with the same reward, relative advantages collapse toward zero and the group supplies no useful learning signal. When the trainable agent's rollouts fail to produce any successful trajectory, one possible solution is to introduce a frozen teacher policy. We refer to the trainable agent as the student. The teacher attempts recovery from a state reached by the student, and its continuation is retained only if the environment verifies success. Yet allowing the teacher to complete the remaining task makes the entire recovered suffix teacher-generated, turning supervision into off-policy imitation and weakening a key benefit of RL: learning on the state distribution induced by the student. We therefore ask: after a teacher has repaired a student failure, when should control return to the student? Our core insight is that the teacher should act only long enough to cross what the student cannot yet handle, then hand back control so that successful experience is again student-generated. We formalize this idea as outcome-verified teacher–student handoff and instantiate it in Verified Handoff for Group Policy Optimization (VH-GPO). Given a verified teacher recovery, VH-GPO returns control to the student at the earliest independently confirmed handoff point found within a bounded search budget. With Qwen2.5-1.5B, VH-GPO increases success rates over GRPO from 59.38% to 94.79% on ALFWorld and from 71.62% to 77.34% on WebShop. Improvements also hold with the 7B student, for which additional teacher and student output tokens amount to 1.77% and 2.61% of regular student rollout output tokens on ALFWorld and WebShop, respectively. Ablation studies validate the design of VH-GPO and demonstrate the effectiveness of its key components. The teacher need not solve every problem for the student; instead, it should provide just enough support when the student struggles to help it find a way forward and gradually learn to solve problems independently. Code is available at https://anonymous.4open.science/r/VO-86EE.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.