HARP: Unifying Reasoning Model and Long-Horizon Agents from Old Teacher Trajectories
Abstract
Reinforcement learning produces experts that solve problems and experts that carry out long-horizon tasks, while deployment calls for one policy that does both. Such consolidation is rarely done once, since a model family spans several sizes and products differ in budget and version. Routes that learn from live teachers, such as on-policy distillation, run the experts again whenever a new student is trained. Yet the experts' own RL training has already logged complete trajectories with their outcomes and behavior probabilities. Such a history records not only what an expert can do, but also how it came to do it. We present HARP, a fully offline framework that consolidates reasoning and long-horizon agency from these histories without running the experts again. The historical path curriculum releases experience along the teachers' own training stages, and the success-ratio budget turns successful trajectories into a provable lower bound on task success. A student then holds a non-decreasing floor on this bound for all acquired tasks. Once frozen, the history supports new students without expert inference or student rollouts for optimization. Experiments with Qwen3.5 at three scales show that HARP is (1) effective, recovering 103.6–105.5% of single-task RL gains on same-size pairs and lifting the 9B student of the 27B history from 22.5% to 39.3% on Terminal-Bench 2.1, (2) reusable, training 4B, 9B, and 27B students from one 27B history with zero new expert inference and 52.3% fewer marginal GPU hours than multi-teacher on-policy distillation at matched quality, and (3) self-sufficient, recovering 97.3% of the weakest-domain gain on eight GPUs with the experts offline. Our code and the record schema of teacher histories are available here.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.