acceptodds
Under review as a conference paper at ICLR 2027

Risk Is Not Reward: Action-Relative Process Rewards for VLM Hallucination Mitigation

Abstract

Vision-language models (VLMs) can assert visual facts unsupported by an image. Decoding-time corrections may repair such errors, but can also damage correct answers. Detecting that the original answer is unreliable is therefore not the same as knowing whether a specified correction will help: queries with identical error risk can have different repair and damage rates. Existing selection strategies based on error risk or visual neglect do not, by themselves, estimate the action-relative quantity needed for intervention. We formulate hallucination mitigation as a one-decision comparison between the corrected continuation and leaving the decoder unchanged. Action-Relative Process Reward (ARPR) gating learns this difference from offline paired outcomes. A process-reward assessor combines static visual evidence with decoder features at the pre-generation boundary, then predicts the fixed action’s net gain without generating either continuation online. A reward gate applies the correction only when its predicted gain warrants it. We derive the optimal decision rule, bound regret from reward-estimation and gating errors, and characterize when additional visual evidence can change the decision. On COCO POPE, accuracy improves by 3.59 percentage points on LLaVA-1.5-7B and 1.01 points on Qwen3.5-9B; applying the correction universally instead reduces LLaVA accuracy to 70.86%.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.