acceptodds
Under review as a conference paper at ICLR 2027

HitchRL: Non-Blocking Teacher Supervision on Idle GPUs in Asynchronous RL

Abstract

Reinforcement learning (RL) with verifiable rewards improves the reasoning of large language models, and group-based methods such as GRPO learn from reward differences among responses to the same prompt. A group whose responses all fail therefore yields no signal, and such groups make up 36.6% of those used for updates in a teacher-free asynchronous RL run on a 1.7B model. A stronger, frozen teacher can still guide them, since it evaluates the student's own responses token by token and gives a direction even when every response fails. Large-scale RL, however, increasingly runs asynchronously on separate GPU pools for generation and optimization, and a teacher does not fit easily into this structure. Given dedicated GPUs, it takes them from the student, and run in step with generation or updates, it makes them wait. We introduce HitchRL, which runs the teacher with no GPUs of its own while no student update waits for an unfinished teacher target. Since a frozen teacher gives a response the same targets whenever it runs, HitchRL defers teacher inference and, with the teacher's weights offloaded to CPU memory, lets it hitchhike on the idle time of the existing GPUs, returning them when generation or optimization needs them, while updates proceed on ready data. With the same eight GPUs, HitchRL raises the eight-benchmark mean accuracy of a 1.7B student from 14.20% with the teacher-free AReaL to 17.04%, the best of all methods compared. This gain over AReaL extends to a larger student, a student from another model family, and code generation with execution feedback. During training, it reaches the same OlympiadBench accuracy as AReaL and as on-policy distillation followed by AReaL in 63% and 76% less wall-clock time, while keeping 0.95 of AReaL's generation throughput with a 30B-A3B teacher and 0.90 even with a 235B teacher.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.