Epistemic Variance Reduction Against Hallucinations in VLM Reinforcement Learning
Abstract
Vision-language models (VLMs) still hallucinate under reinforcement or preference optimization because standard objectives maximize expected reward while ignoring why the reward is uncertain. In ambiguous images, reward models often prefer detailed guesses over calibrated uncertainty, so speculative responses can dominate honest answers in expectation. We formulate this failure mode as a variance-allocation problem: the reward distribution contains aleatoric noise from benign linguistic variability and epistemic uncertainty from conflicting semantic hypotheses about the image. We propose EVR, a variance-aware alignment framework that explicitly decomposes reward variance over semantic response clusters and penalizes the epistemic component during policy optimization. EVR estimates semantic clusters from multiple sampled responses, treats within-cluster reward variance as tolerable noise, and uses between-cluster variance together with a support-weighted local outlier score to discourage low-support factual guesses. We derive a PPO objective and an offline DPO-style variant, and we prove an abstention threshold showing when honest uncertainty should dominate speculative guessing under the shaped reward. The empirical section is organized around standard hallucination benchmarks including POPE, HallusionBench, MMHal-Bench, and M-HalDetect, together with calibration-aware selective answering analysis. This evaluation framing highlights the central design choice of the method: reduce hallucination by targeting epistemic variance rather than indiscriminately penalizing reward noise as a whole.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.