acceptodds
Under review as a conference paper at ICLR 2027

On the Geometry of On-Policy Distillation

Abstract

On-policy distillation (OPD) improves language-model reasoning, but the geometry of its attention weight updates remains poorly understood. We compare OPD with supervised fine-tuning (SFT) and reinforcement learning with verifiable rewards (RLVR). Endpoint diagnostics place OPD in a relaxed off-principal regime: more selective and spectrum-preserving than SFT, but less constrained than RLVR, with qualitative replication on Llama-3.2-3B. Our central trajectory-level finding is subspace locking: cumulative OPD updates align early with a persistent low-dimensional channel. Constraining subsequent training to an early rank-16 subspace leaves OPD performance essentially unchanged, whereas SFT degrades by a much larger margin. Token sparsification and off-policy rollouts preserve the rank trajectory, whereas changes to the objective signal affect its geometry. Mean-broadcast and objective-mixing controls probe this sensitivity further, while exposing confounding by training collapse and unequal gradient scales. Together, these findings characterize OPD's update trajectory and its functional role without positing a universal locking mechanism or claiming efficiency gains.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.