RS-OPSD: Reliable Privileged On-Policy Self-Distillation For Ultra-High-Resolution Remote Sensing VQA
Abstract
Ultra-high-resolution (UHR) remote sensing visual question answering (VQA) requires models to resolve small visual evidence within extremely large images. Existing approaches typically rely on token pruning, visual search, or tool-augmented reasoning at inference time. We instead investigate whether the benefit of *zoom-in visual privilege* can be internalized into the model. We introduce **RS-OPSD**, a reliable privileged on-policy self-distillation (OPSD) framework for UHR remote sensing VQA. To provide high-quality privileged information with explicit question-relevant evidence, we construct **GeoEvidence-6K**, containing 6,750 VQA samples across seven task categories with evidence-region annotations, and develop **Human Feedback-Guided Skill Refinement (HF-SR)** for scalable annotation. To address context loss from tight crops and conflicting signals from imperfect teachers, **RS-OPSD** introduces **Context-Preserving Visual Privilege (CPVP)** and **Correctness-Aligned Distillation (CAD)**. Without any additional visual search or tool calls at inference time, **RS-OPSD** achieves state-of-the-art (SOTA) performance on XLRS-Bench, MME-RealWorld-RS, and LRS-VQA, outperforming previous SOTA models of comparable scale by an average of 4.0 percentage points. Moreover, our 2B variant, **RS-OPD-*Lite***, surpasses most 8B-scale models while achieving the fastest measured inference speed.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.