acceptodds
Under review as a conference paper at ICLR 2027

HT-SCD: Self-Contrastive Decoding via Hallucination Trajectories for Mitigating Object Hallucinations in LVLMs

Abstract

Large vision-language models (LVLMs) demonstrate strong capabilities in image captioning, visual question answering, and multimodal reasoning, yet remain prone to object hallucinations. Existing mitigation methods often require additional training or external detectors, or incur substantial GPU memory overhead. We introduce HT-SCD, a training-free self-contrastive decoding method that uses hallucination trajectories to detect and mitigate object hallucinations online during autoregressive generation, without external models. Whereas conventional contrastive decoding methods construct a generic hallucination-prone negative branch in advance, HT-SCD first identifies specific hallucinations in the generated content and then uses the resulting hallucination trajectories to perform targeted rollback and correction through contrastive decoding. HT-SCD addresses the limited specificity and high computational cost of conventional contrastive decoding, and substantially outperforms existing contrastive decoding methods on MSCOCO-CHAIR and AMBER-G, two benchmarks for hallucination evaluation in open-ended visual generation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.