Commitment Is Not Grounding: Visual-Semantic Rereading for Diffusion MLLMs
Abstract
Diffusion multimodal large language models iteratively resolve answer tokens, yet committing to an answer does not establish visual support for the queried target. We introduce Internal Visual-Semantic Rereading (IVSR), an inference-time controller that audits the native answer at its first valid transfer. A maturity gate selects low-margin commitments for a short visual reread. The reread answer state retrieves an internal visual region, while a contextual question-target representation evaluates its support. A calibrated policy combines this alignment with reread answer margins to keep, reject, or recover the native answer. Backbone weights and prompts remain unchanged; thresholds use labeled held-out calibration. Across DiffusionVL, MMaDA, and LaViDa, IVSR improves all 18 POPE settings by 2.34 F1 points on average and reduces AMBER-D existence hallucination rates by 16.1–47.7%, with 1.4–9.8% measured gated latency overhead. Mechanism diagnostics and ablations support separating inspection priority, answer preference, and target-specific visual evidence in finite-answer verification.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.