MicroCLIP: Connecting Text and Microscopy Images via Contrastive Language–Image Pretraining
Abstract
Understanding fluorescence microscopy images is fundamental to automating biological research. However, existing methods typically learn predefined concepts from closed-set data, limiting transfer across tasks and biological contexts. In this work, we introduce MicroCLIP, a vision–language model that learns generalizable cellular and subcellular representations from natural language supervision for open-set image classification and retrieval. To pretrain MicroCLIP, we construct MicroPool-2M, a dataset of approximately 2.5 million fluorescence microscopy image–text pairs spanning diverse cellular structures, cell lines, and imaging conditions. Aligning microscopy images with biological concepts in natural language is challenging because most images lack descriptive captions and biological knowledge annotations. We therefore progressively establish image–text associations at three levels to provide MicroCLIP with richer semantic supervision: (1) category-level alignment links images to standardized labels defined by subcellular structure–cell-line combinations; (2) knowledge-level alignment incorporates biological priors derived from these labels to capture relationships across species and cell types; and (3) detail-level alignment uses a multimodal large language model to recaption and verify image–text pairs at scale, emphasizing fine-grained visual characteristics. Extensive experiments demonstrate that MicroCLIP achieves leading performance in zero-shot classification and text–image retrieval on microscopy images. It retrieves images with specific cellular characteristics from thousands of candidates and generalizes to unseen biological concepts without additional supervision. These results highlight MicroCLIP’s potential as a foundation model for fluorescence microscopy. Models, training code, and datasets will be publicly released.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.