SCOPD: Sparse-Context On-Policy Self-Distillation for Efficient Vision-Language Models
Abstract
Reasoning vision-language models (VLMs) process images and videos as long sequences of visual tokens, making inference expensive. Training-free token pruning can substantially reduce this cost, but performance degrades sharply under aggressive compression, commonly attributed to irreversible loss of task-relevant visual information. We show that this explanation is incomplete. In a fixed-context Pass@K analysis, we prune each visual representation once and repeatedly sample reasoning trajectories from the same sparse context. Although greedy Pass@1 drops substantially, Pass@K recovers many otherwise failed examples, suggesting that useful visual evidence can remain accessible but is not reliably utilized during reasoning. We call this the representation–utilization gap. Motivated by this observation, we introduce SCOPD, a sparse-context on-policy self-distillation framework in which a student model generates reasoning trajectories from pruned visual tokens while a privileged full-context teacher supervises those on-policy prefixes. SCOPD requires no ground-truth responses, architectural modifications, or inference-time computation. Building on SCOPD, we develop SCOPD+ to focus distillation where visual evidence matters most. A small visual-budget intervention identifies positions most sensitive to additional visual evidence and selectively distills them while backpropagating through only a fraction of responses. At 10% visual-token retention, the Vanilla base model retains only 86.37% of the unpruned model's performance across 13 benchmarks. SCOPD raises this normalized aggregate to 90.49%, while SCOPD+ further improves it to 92.43%. Across token budgets, benchmarks, and pruning operators, our results show that efficient reasoning VLMs depend not only on which visual information survives pruning, but also on how reliably the language model learns to use the sparse representation that remains.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.