OmniMorph: A Channel-Adaptive Unified Generative Model for Cell Painting
Abstract
Cell Painting uses multiplexed fluorescent dyes to label multiple subcellular structures of cells and is a standard assay for large-scale chemical and genetic screening. Due to the vast combinatorial space of chemical and genetic perturbations, using generative models to replace part of these experiments has become a promising avenue. However, the cell lines and perturbations covered by any single dataset are limited, so a model trained on that dataset alone struggles to generalize beyond them. Each dataset is also set up differently depending on its research purpose, such as whether it includes brightfield images or same-batch controls. Therefore, a unified generative model for Cell Painting that supports multiple datasets and multiple tasks is needed. Unlike natural images with a fixed RGB channel space, Cell Painting datasets vary in both the number of channels and the subcellular structures they label, which makes such a unified model difficult to train across datasets. To address this, we propose OmniMorph, to our knowledge the first unified generative model for Cell Painting. OmniMorph is a flow-matching model with a DiT backbone that encodes each fluorescence channel separately. We disentangle the semantics of a channel from its role with two embeddings. The channel embedding marks the subcellular structure the channel represents and is shared across datasets, and the role embedding marks the part it plays in the current task. The model is further conditioned on perturbation embeddings from pretrained encoders and on a natural-language instruction that specifies which task to perform. We train OmniMorph jointly on six datasets and on the four tasks of morphology generation, perturbation prediction, missing-channel completion and brightfield-to-fluorescence translation. It mostly outperforms task-specific methods on image-distribution and biological metrics. Ablations show that the channel embedding is critical and that joint training shares knowledge across datasets and yields gains across tasks, providing a scalable foundation for datasets of heterogeneous channel composition and further generative tasks. Code is available at https://anonymous.4open.science/r/OmniMorph-2EBB/.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.