acceptodds
Under review as a conference paper at ICLR 2027

Self-Steering Multimodal LLMs for Enhanced Knowledge Utilization via Dynamic Information Flow Modulation

Abstract

Multimodal Retrieval-Augmented Generation enhances Multimodal Large Language Models (MLLMs) with external knowledge, yet retrieved evidence is often underutilized during generation. Previous studies typically attribute this issue to insufficient attention allocation or conflicts between contextual and parametric knowledge, overlooking how retrieved information is propagated and preserved in MLLMs. In this work, we revisit this problem from the perspective of internal information flow across Transformer layers. Through layer-wise analyses, we reveal that context-grounded information is better captured in intermediate representations than in final-layer representations, suggesting that ineffective information propagation limits the preservation of retrieved knowledge across layers. To address this, we propose Dynamic Contribution Steering (DCS), a training-free framework that dynamically modulates the contributions of attention, MLP, and residual pathways according to token-level attention distributions and neuron activation statistics. By suppressing irrelevant or dominant internal signals and enhancing context-related representations, DCS improves the utilization of retrieved knowledge without requiring additional training, steering vectors, or extra decoding procedures. Furthermore, we introduce Knowledge Layer Amplification (KLA), which identifies intermediate layers with stronger context-grounded decoding capability and adaptively incorporates their representations during generation. Extensive experiments on MRAG benchmarks across various MLLMs show that the proposed approach improves context-grounded answer generation, establishing information-flow steering as an effective paradigm for improving knowledge utilization in MLLMs.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.