acceptodds
Under review as a conference paper at ICLR 2027

Beyond Confidence: Evidence-Guided Selective Answer Revision for Frozen Vision-Language Model

Abstract

A frozen vision-language model (VLM) may contain a better sampled answer than its first response, yet indiscriminate revision can overwrite a correct baseline. We formulate answer revision as proposal selection followed by baseline-relative acceptance. An evidence-aware candidate ranker (FER) selects a proposal from generation, visual-intervention, and region-conditioned alignment evidence; a Pairwise Utility Gate (UG) estimates replacement utility. We apply a common image-group-separated calibration principle, while the final evaluation construction follows each benchmark’s split structure. The method improves OK-VQA by +1.198 points and VQAv2 by +0.615; on GQA it recovers most of the FER-only degradation and yields a statistically inconclusive +0.064. A matched-data, matched-fold comparison places a compact six-cue schema within one standard error of the expanded FER schema, with small dataset-dependent end-to-end differences. Cross-scale Qwen2.5-VL and model-specific BLIP-2 and LLaVA experiments further show that transfer depends on the backbone and candidate-pool headroom.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.