Seeing What Matters: Visual-Selective On-Policy Self-Distillation for VLM Hallucination Mitigation
Abstract
Vision-language models (VLMs) still suffer from object hallucination in open-ended generation. They often mention objects, attributes, or relations that are not grounded in the image. Preference-based methods such as DPO can reduce hallucinations, but open-ended caption preferences are expensive to collect and often entangle visual faithfulness with style, length, and level of detail. Meanwhile, directly turning object cues into hard pseudo captions can amplify the teacher's linguistic bias and hallucinations. We study how to extend on-policy self-distillation from verifiable language tasks to open-ended multimodal hallucination reduction. Instead of using privileged hints that directly verify an answer, we use training-time object cues to construct a non-verifying privileged visual teacher and distill its soft token distributions on responses sampled by the student itself. We further propose Visual-Selective On-Policy Self-Distillation (VS-OPSD). By contrasting the privileged teacher under the real image and a weak-image condition, VS-OPSD identifies visually sensitive tokens and applies distillation only at those positions. Experiments across two model families and two LLaVA model scales show that VS-OPSD reduces hallucination without preference pairs while largely preserving general multimodal capability. Controlled ablations attribute the gains to concentrating supervision on visually sensitive tokens rather than simply reducing the number of distilled tokens.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.