Partial-Observation-Aware Self-Distillation for Fine-Grained Visual Reasoning
Abstract
Vision-language models often recognize fine-grained details more accurately in a relevant image crop than in the full image. Local-view on-policy self-distillation (OPSD) uses a crop-conditioned teacher to transfer this advantage to a full-image student. However, we uncover an unexpected localization–observation tradeoff: while local-view OPSD strengthens fine-grained observation, it can degrade the ability to localize relevant evidence under full-image inference. Through theoretical analysis and token-level empirical studies, we trace this failure to the partial-observation nature of local-view supervision: it assigns systematically negative optimization signals to tokens that depend on the omitted image context, disproportionately suppressing localization-related tokens. Based on this diagnosis, we propose Partial-Observation-Aware On-Policy Self-Distillation (PA-OPSD), which selectively attenuates unreliable strongly negative optimization signals and introduces a complementary box-guided teacher that restores localization guidance while preserving full-image context. Across fine-grained visual reasoning and localization-oriented benchmarks with 4B and 9B models, PA-OPSD consistently outperforms local-view OPSD and other on-policy baselines, recovering localization while preserving observation benefits. It requires no additional annotations beyond local-view OPSD and no inference-time visual operations.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.