Learn in Space: On-Policy Distillation via Spatial Divergence for Vision-Centric Reasoning
Abstract
Spatial awareness matters in vision-centric reasoning. Despite the effectiveness of its fine-grained supervision, on-policy distillation (OPD) remains spatially agnostic during optimization. Fundamentally, the same visual evidence can admit many spatially equivalent region descriptions, consequently leading to a vast number of distinct token serializations. Token-level KL treats these serializations as distinct outcomes in vocabulary space, which provides biased learning signals for spatially valid grounding behavior, ultimately limiting performance. To bridge these gaps, we introduce MixOPD, which uses Spatial Divergence to measure the geometric discrepancy between each on-policy student grounding proposal and corresponding teacher-supported evidence distribution estimated via Teacher-Forced Constrained Search, thereby spatially calibrating the token-level KL objective. MixOPD integrates seamlessly into typical OPD frameworks and delivers theoretically grounded performance gains. Extensive experiments demonstrate that MixOPD consistently improves both standard and entropy-aware OPD. It achieves highly competitive performance against representative post-training baselines and recent alternatives, while presenting favorable inference latency and stable training dynamics.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.