acceptodds
Under review as a conference paper at ICLR 2027

Partial-Observation-Aware Self-Distillation for Fine-Grained Visual Reasoning

Abstract

Vision-language models often recognize fine-grained details more accurately in a relevant image crop than in the full image. Local-view on-policy self-distillation (OPSD) uses a crop-conditioned teacher to transfer this advantage to a full-image student. However, we uncover an unexpected localization–observation tradeoff: while local-view OPSD strengthens fine-grained observation, it can degrade the ability to localize relevant evidence under full-image inference. Through theoretical analysis and token-level empirical studies, we trace this failure to the partial-observation nature of local-view supervision: it assigns systematically negative optimization signals to tokens that depend on the omitted image context, disproportionately suppressing localization-related tokens. Based on this diagnosis, we propose Partial-Observation-Aware On-Policy Self-Distillation (PA-OPSD), which selectively attenuates unreliable strongly negative optimization signals and introduces a complementary box-guided teacher that restores localization guidance while preserving full-image context. Across fine-grained visual reasoning and localization-oriented benchmarks with 4B and 9B models, PA-OPSD consistently outperforms local-view OPSD and other on-policy baselines, recovering localization while preserving observation benefits. It requires no additional annotations beyond local-view OPSD and no inference-time visual operations.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.