acceptodds
Under review as a conference paper at ICLR 2027

On the Orthogonal Geometry of On-Policy Distillation

Abstract

On-policy distillation (OPD) trains a student on its own sampled trajectories using dense token-level teacher feedback. Standard low-rank adaptation (LoRA) makes OPD more parameter-efficient, but restricts learning to an initialization-dependent channel whose suitability for OPD is unclear. We introduce OP-OPD and OM-OPD, two orthonormal spectral initialization methods for low-rank OPD. They initialize the LoRA row space with, respectively, the principal and minor right-singular directions of pretrained weights, while a zero output factor preserves the initial model function. A first-step analysis shows that this choice projects the OPD gradient onto a selected subspace, making the initialization channel explicit. On three mathematical reasoning benchmarks, both methods improve over standard LoRA-OPD on aggregate mean@16 and maj@16 in two teacher–student settings. Update-trajectory and spectral diagnostics show that the two variants follow distinct pretrained channels, consistent with the proposed routing interpretation. The results identify initialization geometry as a design dimension for OPD, with gains observed across both evaluated teacher–student pairs.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.