Certify Before You Intervene: Pathwise Intervention-Certified Decoding for Multimodal Generation
Abstract
Vision-language models (VLMs) describe objects, attributes and findings that the image does not show. Inference-time corrections, such as retrieving evidence and regenerating part of the response, reduce these hallucinations on average, but they also turn correct responses into wrong ones and delete useful content. Whether a correction helps depends on the specific revision, which uncertainty and evidence relevance measure only indirectly. We introduce Pathwise Intervention-Certified Decoding (PICD), which revises a response only when a calibrated lower bound on the benefit of that revision is positive. A critic, trained on paired continuations that share an input and a prefix, predicts how a revision changes the hallucination loss and the utility of the completed response. PICD calibrates the largest overestimation of this critic along each response, so that its bounds hold for whichever candidate the decoder selects, and recalibrates every round on the states that earlier rounds produce. Under exchangeability, with a probability the operator chooses, every accepted revision on a response's path lowers hallucination loss and stays within a utility tolerance. Across three VLMs and five general-domain and medical benchmarks, the critic predicts which revisions lower hallucination loss better than uncertainty and relevance signals, and PICD keeps the share of responses with a violating revision below each requested rate we test, at the cost of revising fewer responses than an uncalibrated critic.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.