Beyond the Current Prediction: Tracking Decoding Trajectories for Object Hallucination Detection in LVLMs
Abstract
Large vision-language models (LVLMs) can confidently generate objects that are plausible under the generated context but absent from the image. Existing training-free object hallucination detectors mainly score each generated object using signals available at its current generation step. However, autoregressive generation can obscure uncertainty revealed earlier in the decoding process. At each step, the model assigns a probability to the generated token, but once the token is selected, this probability is no longer retained in subsequent decoding. Later predictions instead condition on the selected token itself. Consequently, a token generated with low confidence can shape the subsequent context and make a later hallucinated object appear locally confident despite uncertainty observed earlier in the generation trajectory. Motivated by this observation, we ask whether retaining uncertainty signals from preceding decoding steps can improve object hallucination detection. We introduce History-Aware Risk Tracking (HART), a simple training-free method that estimates hallucination risk using both the current object prediction and uncertainty observed throughout the preceding generation process. Specifically, HART considers both the current object prediction and uncertainty accumulated over preceding decoding steps, with greater weight assigned to more recent steps. We use negative log-likelihood as a simple token-level measure of prediction uncertainty. This formulation retains information that would otherwise no longer be directly available as generation proceeds, without requiring additional training or auxiliary generation. Experiments across multiple LVLMs on MSCOCO and Pascal VOC show that incorporating decoding history improves over current-step scoring and achieves strong object hallucination detection performance compared with existing training-free methods. Our results show that uncertainty from preceding decoding steps provides a complementary signal for object hallucination detection beyond the current object prediction alone.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.