Blind at Gigapixel Scale: Visual Information Utilization in MLLMs for Whole-Slide Images
Abstract
Pathology multimodal large language models (PMLLMs) have attracted increasing attention for their ability to process whole-slide images (WSIs) and support tasks such as diagnosis assistance, question answering, and report generation. Despite strong performance on existing benchmarks, whether these models effectively utilize slide-specific visual evidence remains unclear. Unlike normal images, WSIs require highly compressed visual representations due to their gigapixel scale, raising questions about whether diagnostically relevant information can be accessed and incorporated into model decisions. We systematically analyze visual evidence utilization in PMLLMs by examining their inputs, representations, and predictions. We evaluate representative models across different input paradigms under controlled visual interventions and introduce VAPath, a WSI–QA–ROI benchmark with pathologist-annotated visual anchors for assessing whether model-derived visual signals align with pathology-relevant regions. Our analysis reveals a gap between visual information availability and decision-level utilization: visual interventions often have limited effects on model predictions, while an independent pathology retrieval procedure identifies stronger ROI-associated patterns in WSI patch features than model-derived signals. We further show that multimodal instruction tuning can improve answer fitting without increasing the measured marginal contribution of WSI visual information under our evaluation protocol. These findings demonstrate that benchmark accuracy alone provides an incomplete measure of WSI understanding and motivate evaluations that directly characterize how visual evidence contributes to PMLLM decisions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.