acceptodds
Under review as a conference paper at ICLR 2027

Beyond Average Gains: Replay-Based Post-Intervention Control

Abstract

Training-free mitigation methods can reduce multimodal hallucinations on aver- age, but fixed, always-on use still leaves a practical problem: an intervention can improve aggregate performance while turning some originally correct outputs into wrong ones. We study this as post-intervention risk control, where the goal is to identify harmful method-induced changes after generation and apply the small- est possible correction. We find that standard generation-time signals, including confidence, margin, entropy, selective prediction, and self-consistency, provide little information about whether an intervention was warranted. In contrast, re- playing the committed intervention output through the unmodified frozen back- bone exposes a substantially stronger signal for separating harmful from helpful changes. Based on this observation, we propose RAPIC, a lightweight replay- based control layer that selectively retains or reverts intervention outputs. In our primary LLaVA-1.5 + VGA setting, RAPIC improves accuracy by +1.21, +1.61, and +0.67 points on MSCOCO, AOKVQA, and GQA, respectively, with paired image-clustered bootstrap confidence intervals excluding zero. Repeated calibra- tion shows that changed-answer labeling is more efficient than uniform sampling.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.