MisLocus: A single-cell imaging benchmark for protein mislocalization
Abstract
Genetic variants that alter the location of proteins within cells can disrupt protein function and contribute to disease. Microscopy can detect protein mislocalization, but benchmarking image representations to automate this task requires data for reference (healthy) proteins plus protein variants, with a replicate structure that separates biological signal from technical variation. We introduce MisLocus, a single-cell imaging benchmark for variant-induced protein mislocalization. It contains ∼3.3 million single-cell images, covering 250 reference proteins and 1,291 variants, with experimental replicates spanning imaging wells, plates, and batches. We evaluate nine image representations from CellProfiler, Cytoself, SubCell, and MorphEm, including handcrafted features, pretrained learned representations, and models retrained or fine-tuned on MisLocus. Representations are assessed along three complementary axes: detecting variant-induced mislocalizations, predicting each variant’s clinical pathogenicity based on mislocalization phenotype, and discriminating annotated reference protein localization patterns. We find no single representation dominates all evaluation axes. Engineered features, paired with supervised models trained to predict localization shifts, detect variant-induced mislocalization competitively with the best learned representations, whereas learned representations better organize proteins according to shared localization patterns. MisLocus establishes a controlled and reproducible framework for benchmarking single-cell microscopy representations on protein mislocalization phenotyping.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.