acceptodds
Under review as a conference paper at ICLR 2027

Mitigating Hallucination in Visual Language Models with Partial Transport KL-constrained Decoding

Abstract

Vision-language models frequently describe objects that are not present in the image, limiting their use in settings where faithfulness matters. Existing inference-time mitigation methods typically rely on contrastive decoding, attention manipulation, or representation steering, but suppressing hallucinations can also remove valid visual content. We introduce POTION (Partial Optimal Transport Inference with KL ProjectiON), a training-free inference-time method that steers autoregressive VLM decoding toward visually grounded outputs. At each generation step, POTION solves a null-penalized partial optimal transport problem between top- candidate-token embeddings and cached image-patch hidden states. The partial formulation automatically routes visual support mass to grounded candidates while assigning unsupported candidates to a null sink, yielding per-candidate evidence and risk scores that produce a logit correction. This correction is then projected onto a KL-ball centered at the base model's distribution, preventing aggressive hallucination suppression from collapsing object coverage. Experiments on MSCOCO and AMBER show consistent reductions in object hallucination across Qwen2.5-VL and Llama-3.2-Vision. POTION achieves state-of-the-art performance in mitigating object hallucinations, substantially reducing CHAIR while maintaining comparable object recall.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.