Global Relevance Is Not Enough: Representative Candidate Construction and Residual-Based Evidence Preservation for Efficient, Reliable Medical MLLM Inference
Abstract
Medical Multimodal Large Language Models (MLLMs) have shown great potential in clinical visual question answering and medical image interpretation, but their high inference cost limits efficient deployment in real-world clinical settings. Moreover, in safety-critical medical scenarios, pursuing efficiency alone is far from sufficient. We observe that existing acceleration methods, although able to preserve aggregate performance, can still induce substantial instance-level prediction shifts. To address this issue, we propose MedR2R (Representative-to-Residual Compression), a representative candidate construction and residual-based evidence preservation framework for efficient and reliable medical MLLM inference. Rather than determining token importance from global relevance alone, MedR2R adopts a two-stage coarse-to-fine selection strategy. First, after the vision encoder, it constructs a semantically representative candidate set under a relaxed intermediate token budget. Specifically, MedR2R models pairwise visual-token relationships using cosine similarity and combines attention-derived global relevance with marginal representation gain, encouraging the candidate set to represent the overall visual space rather than only highly attended regions. Second, after the visual tokens are contextualized by an intermediate language-model layer, MedR2R further refines the candidate set according to normalized residuals. Tokens that can be accurately approximated from a compact basis are treated as removable, whereas tokens with large normalized residuals are preserved as hard-to-replace visual evidence. Experiments across multiple medical VQA datasets and MLLMs demonstrate that, even under substantially reduced visual-token budgets, this coarse-to-fine two-stage selection strategy maintains competitive task performance while effectively mitigating harmful instance-level prediction shifts. These results further indicate that medical MLLM acceleration should not focus solely on aggregate performance, but should also account for the stability of individual predictions. Our code is available at: https://anonymous.4open.science/r/MedR2R.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.