MedScissor: Towards General and Efficient Medical Image Understanding with Vision-Language Models
Abstract
Medical vision-language models have shown substantial promise across various medical tasks. However, the large number of visual tokens generated from medical images introduces significant computational and inference overhead. Existing approaches either rely on internal attention signals, complicating their direct compatibility with efficient attention mechanisms such as FlashAttention, or are primarily designed for 3D medical imaging, limiting their applicability to broader medical imaging scenarios. To address these limitations, we propose MedScissor, a general and efficient visual token compression framework for medical vision-language models. Specifically, MedScissor first combines the semantic similarity score and neighborhood consistency score of medical visual tokens to identify redundant relationships among tokens, and then progressively merges related groups until the target compression budget is reached. Finally, MedScissor jointly considers local centrality score and global deviation score to retain one representative token with complementary information from each group. Extensive experiments demonstrate the effectiveness of MedScissor across diverse scenarios. On six representative 2D medical imaging benchmarks, MedScissor reduces the number of visual tokens by 50% while retaining 98.6% of the original performance, achieving a 1.4× prefill speedup and a 49.2% reduction in KV cache usage. On two representative 3D medical imaging benchmarks, MedScissor reduces the number of visual tokens by 70% while retaining 99.0% of the original performance, achieving a 2.1× prefill speedup and a 65.6% reduction in KV cache usage. Our code is available in the supplementary materials, and all data will be released on GitHub.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.