Know More but Say Less: Bridging the Generation–Discrimination Gap of VLMs via Latent-Guided Verification
Abstract
Vision-Language Models (VLMs) can turn unreliable visual inputs into fluent descriptions that downstream systems accept as facts, compromising reasoning and decisions. We attribute this failure to the generation-discrimination (G-D) gap, in which unreliability is omitted during descriptive generation but recognized under explicit verification. Yet its mechanisms remain unclear, and conventional hallucination mitigation through direct internal modification is ill-suited to addressing it. We therefore analyze hidden activations, visual attention, and output distributions. Unreliability evidence remains linearly separable in intermediate representations and visual tokens receive substantial attention, but this distinction is weakly reflected in output probabilities and generated text. Based on these findings, we propose Bridge, which estimates unreliability from prefix activations and selectively injects hierarchical verification prompts without updating VLM parameters or directly modifying internal states. Across four benchmarks, Bridge outperforms all VLM-based baselines on the primary reliability metric. Relative to descriptive generation, it improves F1 by up to 0.845 and Recall by 0.550. On our self-designed 600-sample evaluation across three tasks, it achieves 84.2% task success and reduces unreliable information propagation by 11.7 percentage points, balancing reliability and utility. Our implementation is available at https://anonymous.4open.science/r/GDBRIDGE-F7CD.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.