acceptodds
Under review as a conference paper at ICLR 2027

HARP: Unifying Reasoning Model and Long-Horizon Agents from Old Teacher Trajectories

Abstract

Reinforcement learning produces experts that solve problems and experts that carry out long-horizon tasks, while deployment calls for one policy that does both. Such consolidation is rarely done once, since a model family spans several sizes and products differ in budget and version. Routes that learn from live teachers, such as on-policy distillation, run the experts again whenever a new student is trained. Yet the experts' own RL training has already logged complete trajectories with their outcomes and behavior probabilities. Such a history records not only what an expert can do, but also how it came to do it. We present HARP, a fully offline framework that consolidates reasoning and long-horizon agency from these histories without running the experts again. The historical path curriculum releases experience along the teachers' own training stages, and the success-ratio budget turns successful trajectories into a provable lower bound on task success. A student then holds a non-decreasing floor on this bound for all acquired tasks. Once frozen, the history supports new students without expert inference or student rollouts for optimization. Experiments with Qwen3.5 at three scales show that HARP is (1) effective, recovering 103.6–105.5% of single-task RL gains on same-size pairs and lifting the 9B student of the 27B history from 22.5% to 39.3% on Terminal-Bench 2.1, (2) reusable, training 4B, 9B, and 27B students from one 27B history with zero new expert inference and 52.3% fewer marginal GPU hours than multi-teacher on-policy distillation at matched quality, and (3) self-sufficient, recovering 97.3% of the weakest-domain gain on eight GPUs with the experts offline. Our code and the record schema of teacher histories are available here.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.