Content-Bounded Phase Consensus: Mitigating Position-Driven Source Preference in Multi-Image Vision-Language Models
Abstract
Multi-image vision–language models (VLMs) must associate visual evidence with its source, yet their predictions can depend on source position. We study this behavior through Boundary Keys, the Key representations of model-native tokens delimiting visual sources. Across representative models from Qwen3-VL, InternVL3.5, and Molmo2-O, identical visual inputs exhibit position-dependent source attention, while early-layer Boundary Keys contain a dominant source-shared component. RoPE maps this shared component to source-dependent realizations, and counterfactual phase reassignment changes boundary source preference even when native source-specific residuals are retained. Motivated by this analysis, we propose Content-Bounded Phase Consensus (CBPC), a lightweight training-free intervention that selectively contracts the RoPE-induced phase deviation while preserving the cross-source mean. CBPC derives an input- and layer-dependent correction strength from native visual-key-centroid statistics without additional parameters, answer labels, or strength search. Across Mantis, MuirBench, MIRB, and QBench2, CBPC yields consistent gains, improves correctness under meaning-preserving image reorderings, and incurs modest LM forward overhead with negligible additional peak memory.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.