acceptodds
Under review as a conference paper at ICLR 2027

Rollout-Visual-Commit: Mitigating Object Hallucination via Commit-Time Multimodal Scoring

Abstract

Large vision–language models (LVLMs) answer complex visual questions fluently, yet still invent absent objects, deny present ones, or extend captions after an unsupported token. Fine-tuning remedies require extra supervision, current-step interventions act before a candidate has served as a query, and trajectory or reward methods add latency. All leave the multimodal state induced by a provisional next token unused. We propose Rollout-Visual-Commit (RVC), a decoding rule without LVLM post-training. RVC rolls top-K candidates forward in one branch-isolated packed pass, scores each candidate-induced state by probe-selected visual-attention mass times next-state logit margin, adds the standardized response to its base log-probability, and keeps only the winning KV entry. Thus, an unsupported but fluent token can be rejected before entering the prefix. On nine POPE settings with LLaVA-1.5 and Qwen2.5-VL, RVC attains the best Accuracy and F1 among the included methods (average per-setting margins 2.57/2.66 and 1.44/2.01). It also leads every reported CHAIR and AMBER metric, including AMBER coverage.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.