acceptodds
Under review as a conference paper at ICLR 2027

G²-OPD: Grounded and Graspable On-Policy Distillation for Dynamic Spatial Reasoning

Abstract

Compact vision-language models are attractive for embodied applications, yet transferring dynamic spatial reasoning from stronger teachers remains challenging. Fixed-trace distillation suffers from a mismatch between teacher-provided reasoning prefixes and the student's own inference trajectories, while outcome-based reinforcement learning provides sparse corrective signals when successful rollouts are rare. We therefore study on-policy distillation (OPD), which supervises student-generated trajectories with dense teacher distributions. However, standard OPD treats reasoning positions uniformly, ignoring whether a prediction is grounded in spatiotemporal evidence and whether the corresponding correction is learnable by the student. To this end, we introduce G-OPD, a Grounded and Graspable OPD framework that adaptively allocates teacher supervision along the student's trajectory. Grounding uses counterfactual video probes to measure sensitivity to visual content, temporal order, and motion. Graspability captures corrections that are both needed and learnable by combining teacher–student disagreement with teacher support for the student's leading candidates. We then combine these two signals through a regularized supervision-allocation rule that redistributes full-vocabulary reverse-KL supervision while preventing excessive concentration. These signals reweight full-vocabulary reverse-KL supervision without task-specific process rewards or additional inference-time modules. Ultimately, G-OPD achieves 51.38% accuracy on VLM4D and 34.49% accuracy on DSR-Bench, improving over baseline by 11.02 and 12.12 percentage points, respectively.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.