acceptodds
Under review as a conference paper at ICLR 2027

ViECo: Visual Energy-Conservation Decoding for Hallucination Mitigation in Large Vision-Language Models

Abstract

Large vision-language models (VLMs) often generate visually unsupported yet linguistically plausible claims, a failure mode known as hallucination. Existing training-free mitigation methods contrast image-conditioned logits with distorted-image or null-image branches, or steer latent states toward visual features, but rarely ask a more basic question: where does the probability of a visual claim come from? Inspired by recent findings that language-model logits exhibit low-rank structure and that hallucinations leave energy inconsistencies along decoding trajectories, we propose ViECo, a training-free framework that views VLM hallucination as visual energy non-conservation. ViECo decomposes image-conditioned logits into a low-rank language-prior subspace and an orthogonal visual-residual component, then penalizes visual claims whose probability is dominated by prior projection with weak residual support and temporal energy spill. The method requires no parameter updates, external detectors, or image captions. ViECo achieves 84.40 POPE-Adversarial F1 (+7.95 over vanilla decoding), reduces CHAIRs from 46.60 to 30.60, and obtains the highest MMHal score (3.12) among all training-free methods, with only 1.45x overhead and the lowest prompt-variation sensitivity.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.