acceptodds
Under review as a conference paper at ICLR 2027

U-TRACE: Uncertainty-Guided Adaptive Evidence Acquisition for Open-Domain VQA

Abstract

Open-domain visual question answering requires systems to answer questions that may involve visual understanding, external knowledge, or multimodal reasoning. Existing approaches typically rely on fixed pipelines and single-pass generation, such as retrieval or visual grounding, without explicitly assessing whether the current answer is reliable or which intervention would improve it. To address these limitations, we introduce U-TRACE, an uncertainty-guided framework for adaptive evidence acquisition in open-domain VQA. U-TRACE probes the vision–language model under complementary visual and textual contexts to construct a model-specific behavioral state that captures semantic dispersion, context sensitivity and prediction stability. ProbeDetector uses this state to estimate the reliability deficit of the initial prediction and determine whether further evidence acquisition is required. For selected instances, a history-conditioned agent chooses the next evidence-seeking action and specifies its request when necessary. The resulting report is added to the evidence history and used to update the uncertainty state, so that each subsequent decision reflects what has already been observed. To learn the adaptive policy, we propose execution-verified learning, which compares alternative next actions from the same evidence history and assigns supervision based on their impact on final-answer quality. Experiments on OK-VQA and A-OKVQA show that U-TRACE consistently improves its base VLMs and outperforms existing open-domain VQA methods.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.