acceptodds
Under review as a conference paper at ICLR 2027

Learn in Space: On-Policy Distillation via Spatial Divergence for Vision-Centric Reasoning

Abstract

Spatial awareness matters in vision-centric reasoning. Despite the effectiveness of its fine-grained supervision, on-policy distillation (OPD) remains spatially agnostic during optimization. Fundamentally, the same visual evidence can admit many spatially equivalent region descriptions, consequently leading to a vast number of distinct token serializations. Token-level KL treats these serializations as distinct outcomes in vocabulary space, which provides biased learning signals for spatially valid grounding behavior, ultimately limiting performance. To bridge these gaps, we introduce MixOPD, which uses Spatial Divergence to measure the geometric discrepancy between each on-policy student grounding proposal and corresponding teacher-supported evidence distribution estimated via Teacher-Forced Constrained Search, thereby spatially calibrating the token-level KL objective. MixOPD integrates seamlessly into typical OPD frameworks and delivers theoretically grounded performance gains. Extensive experiments demonstrate that MixOPD consistently improves both standard and entropy-aware OPD. It achieves highly competitive performance against representative post-training baselines and recent alternatives, while presenting favorable inference latency and stable training dynamics.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.