acceptodds
Under review as a conference paper at ICLR 2027

See What I See: On-Policy Distillation of Visual Evidence, Not Just Words

Abstract

On-policy distillation (OPD) has become a core technique in the post-training of large language models and is increasingly applied to multimodal ones. It trains a student with dense token-level feedback from a teacher on the student's own rollouts. However, this feedback supervises what the student says, not what it sees: even a correct token may draw little from the image, or draw from the wrong region. We find that during OPD the student's output distribution converges to the teacher's, while an evidence gap persists. The gap has two sides: (i) a few load-bearing tokens carry most of the teacher's evidence, yet their supervision is diluted across the whole response; (ii) on these tokens the student's evidence misses the ground-truth region, even where its next-token distribution already matches the teacher's. To close this gap, we propose SWIS ( See What I See ), which distills the teacher's visual evidence directly rather than through its next-token distribution alone. SWIS models evidence as a joint distribution over token–patch pairs and aligns the student's with the teacher's, matching both which tokens carry the evidence and where on the image it lands. Across ten benchmarks, SWIS improves average accuracy by 1.8 pp, gaining most where answers depend on the image, and closes the evidence gap under the same readout that exposed it.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.