acceptodds
Under review as a conference paper at ICLR 2027

Quotient-Set Evidence Fusion for Robust Multimodal Reasoning

Abstract

Reliable reasoning over multiple images and document pages requires combining distinct facts, yet repeated content can disproportionately influence predictions even when it adds no information. We make multimodal reasoning resistant to artificial repetition while preserving the complementary and conflicting evidence needed to answer a question. We introduce Quotient-Set Evidence Fusion (QSEF), which identifies redundancy relative to the question and adjusts each evidence unit's attention weight according to its estimated repetition count. QSEF corrects attention before pooling independently encoded inputs into a compact evidence representation, separating the relevance of information from its frequency while retaining every input. Across eligible subsets of six multi-image benchmarks, including MuirBench and BLINK, and three vision–language backbones, QSEF improves accuracy-based scores by 1.4–2.7 percentage points over the strongest compared baseline in all 18 comparisons. Our analysis establishes exact repetition invariance for correctly identified copies with matching representations and bounds representation drift under imperfect redundancy estimates, clarifying when robust multimodal evidence aggregation is achievable.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.