acceptodds
Under review as a conference paper at ICLR 2027

GroundPO: Policy Optimization via Internalization and Distillation for Visual Grounding

Abstract

Visual grounding becomes substantially more challenging when a query identifies its target through implicit comparison, spatial relations, and other forms of reasoning. Recent advances in on-policy distillation (OPD) provide a promising paradigm for such reasoning-intensive grounding. However, we find that distilling from a teacher with privileged information can cause training collapse, manifested as Privilege-Induced Behavioral Shift (PrivShift): privilege-specific artifacts and excessively long reasoning trajectories are inherited by the student. To address this issue, we propose GroundPO, a two-stage framework following the principle of internalizing the knowledge before teaching the policy. We first perform Policy Internalization to absorb privileged information into the teacher policy, encouraging reasoning that is free from privilege leakage and maintains an appropriate response length. Policy Distillation then uses the resulting internalized policy as the teacher to transfer its grounding capability through OPD to a student without privileged input. We further introduce DeepGround, comprising a high-quality training set and a challenging benchmark spanning five tracks of reasoning-intensive grounding. Experiments show that GroundPO substantially mitigates PrivShift, while GroundPO-27B achieves state-of-the-art performance, surpassing Gemini-3.1-Pro. Further analysis shows that GroundPO improves fine-grained perception and relative-position reasoning, demonstrating benefits beyond visual grounding on V* Bench.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.