GAZEBREAST-FM: AN OMNI BREAST IMAGING FOUNDATION MODEL WITH GAZE-GUIDED DIAGNOSTIC EVIDENCE ALLOCATION
Abstract
Breast cancer diagnosis relies on complementary information from different breast imaging modalities, yet developing separate models for individual modalities and tasks requires substantial manual annotation and limits representation sharing across imaging domains. Foundation models can reduce this dependence by learning transferable representations through large-scale pre-training, but existing foundation models for breast imaging remain largely modality-specific and are not designed to learn jointly from heterogeneous breast imaging data within a unified framework. GazeBreast-FM is an Omni breast imaging foundation model pre-trained on 302,047 images from 23 datasets. The framework jointly optimizes MAE-style masked reconstruction and CLIP-style image–report alignment, while introducing a spatial attention prior generated from simulated radiologist-like gaze trajectories to allocate sparse diagnostic evidence between masked reconstruction and visible-region semantic learning. A salient-evidence preservation constraint prevents diagnostically relevant information from being completely masked, and diagnostic-evidence coverage is used to dynamically coordinate the two learning objectives. Diagnostic evidence that remains visible after masking is further used for clinical concept learning, while a report-derived Clinical Graph captures continuous clinical relationships across cases and provides structured semantic supervision. Across multiple downstream tasks and breast imaging modalities, GazeBreast-FM achieves state-of-the-art performance, demonstrating transferable representations that support global diagnosis, lesion-level spatial understanding, and fine-grained clinical semantics.https://github.com/aaaaaaa12344/GazeBreast-FM
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.