DriftDPO: Mechanism-Guided Preference Optimization for Hallucination Mitigation in Deep-Fusion Vision-Language Models
Abstract
Object hallucination in vision-language models (VLMs) is widely attributed to language-prior dominance, motivating contrastive decoding (CD) methods that suppress this prior. However, CD consistently worsens hallucination on modern deep-fusion VLMs, suggesting the prevailing theory does not hold for these architectures. Through token-level probing of internal representations, we discover that the theory is not merely incomplete but inverted: on deep-fusion VLMs, hallucination is driven by visual-mean alignment rather than language-prior dominance, and this mechanism is architecture-gated—active in deep-fusion models, absent in shallow-projection models, and reversed in next-generation ones. Based on this finding, we propose DriftDPO, which derives its preference signal entirely from the model's own hidden states, requiring no external annotation. A per-token analysis reveals that the caption-level signal is biased by function-word dominance, necessitating a direction correction step to align the training objective with the object-level mechanism. DriftDPO achieves state-of-the-art hallucination reduction on multiple benchmarks while increasing coverage at zero inference overhead—the only method to simultaneously achieve all three. Mechanism verification confirms that training genuinely internalizes the signal into model weights rather than superficially suppressing it, revealing a coupling property unique to self-supervised rewards.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.