EvidenceFlow: Repaying Visual Evidence Debt for Test-Time Multimodal Reasoning
Abstract
Test-time scaling can improve multimodal reasoning without updating model parameters, but longer reasoning may not provide the visual evidence required by later steps. Existing visual refocusing methods replay image tokens based on their relevance to the current reasoning state, without examining whether realized visual support matches that state. We study the resulting mismatch between the current visual-demand distribution and the realized visual-support distribution. We define this mismatch as Visual Evidence Debt, quantified as one minus the overlapping probability mass of the two distributions. Its equivalent positive-residual form localizes visual tokens whose demand share exceeds their realized-support share. We therefore propose EvidenceFlow, a training-free framework that performs marginal debt repayment under a fixed replay budget via a saturating monotone submodular objective. Experiments and controlled analyses across multiple MLLMs and benchmarks show improved reasoning accuracy over standard thinking and a more favorable accuracy–compute trade-off than the evaluated replay baselines. Our code is available in the supplementary material.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.