acceptodds
Under review as a conference paper at ICLR 2027

HARP: On-Policy Distillation of Hierarchically Routed Planning, Execution, and Revision Experts

Abstract

Long-horizon tool-use agents remain difficult to train because task-level outcome rewards provide sparse credit across trajectories involving many interdependent decisions. Process-level supervision alleviates this sparsity with finer-grained feedback, but typically treats intermediate errors through a largely homogeneous signal. In interactive environments, however, failures are structurally heterogeneous: an agent may pursue an incorrect subgoal, execute a valid plan with an incorrect tool call, or fail to diagnose and recover from an earlier mistake. Effective supervision should therefore capture not only where credit should be assigned, but also what kind of failure occurred and how it should be corrected. We introduce HARP (Hierarchically Arranged Role Policies), a training framework that decomposes failures into planning, execution, and recovery errors and assigns them to role-specialized Planner, Executor, and Repairer components coordinated by verification and routing mechanisms. HARP uses this structured system as a training-time teacher and distills its role-specific corrective supervision on-policy into a single base model, avoiding the full scaffold at deployment. On AppWorld with Qwen3-1.7B, 4B, and 8B, the distilled HARP model improves task goal completion over the base model at every size and outperforms GRPO at 1.7B and 4B, while GRPO remains stronger at 8B. Component ablations and routing diagnostics further characterize the contribution of each role.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.