Schrödinger's Prior: Quantum-inspired Representation Intervention for Hallucination Mitigation in LVLMs
Abstract
Large Vision-Language Models (LVLMs) achieve strong multimodal performance but remain vulnerable to hallucinations, where language priors can override visually grounded evidence. Existing mitigation methods therefore tend to suppress prior-related information, implicitly treating priors as intrinsically harmful. We challenge this assumption and reveal a context-conditioned prior duality: the same prior can provide useful knowledge or induce hallucination depending on the current image, prompt, and decoding state. We further show that, when beneficial and harmful prior effects coexist, no fixed context-independent attenuation can improve both. Based on this finding, we propose SchröHallu, a training-free, quantum-inspired inference framework that models unresolved prior utility as a superposition of hallucination-inducing and knowledge-beneficial states. SchröHallu strengthens visual evidence, decomposes hidden representations into evidence, prior, and residual components, and measures whether contextual evidence sufficiently resolves the prior toward a harmful state. Only then is a minimum-norm correction applied along the decision-relevant direction of the prior subspace, preserving grounded evidence and residual capabilities. Extensive experiments across three representative LVLMs consistently reduce hallucinations while largely preserving perceptual and reasoning performance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.