acceptodds
Under review as a conference paper at ICLR 2027

When Delimiters Mislead: Visual-Text Regrounding for Multi-Image Hallucination Mitigation

Abstract

Multimodal large language models (MLLMs) are increasingly capable of reasoning over multiple images, but this capability also introduces new hallucination risks. In this paper, we reveal that delimiter tokens designed to separate multiple images are associated with two distinct internal patterns on the visual and textual sides: Visual-Delimiter Leakage and Text-Delimiter Sink, both of which contribute to multi-image hallucination. To address these issues, we propose a training-free method called Delimiter-Aware Regrounding (DAR). DAR consists of two modules corresponding to the two internal patterns: Delimiter Ownership Correction (DOC) mitigates Visual-Delimiter Leakage by restoring visual-token ownership, while Visual Evidence Redirection (VER) strengthens image-specific visual grounding to alleviate Text-Delimiter Sink. Extensive experiments on multi-image benchmarks across multiple MLLMs show that DAR effectively enhances multi-image reasoning while preserving the structural role of delimiter tokens in separating visual inputs. Moreover, DAR provides a new perspective on hallucination mitigation in broader multi-input scenarios, such as video reasoning.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.