acceptodds
Under review as a conference paper at ICLR 2027

scMOD: Completing and Generating the Virtual Cell with Trimodal Latent Diffusion

Abstract

Virtual cells aim to simulate cellular states, but a faithful model must capture multiple molecular modalities and operate when any subset is missing. Existing translation methods are tied to predefined assay pairs and require an observed source. Shared-latent VAEs provide broader coverage, but aligning heterogeneous assays within one latent space can obscure modality-specific information, while their simple priors limit de novo generation. We introduce scMOD, to our knowledge the first latent diffusion model for trimodal RNA, chromatin accessibility (ATAC), and surface protein (ADT) data. scMOD combines assay-specific variational codecs with a shared diffusion Transformer trained by block-causal flow matching over randomized modality orders. The first block learns marginal generation, and subsequent blocks learn conditional generation from a clean multimodal prefix. This design allows a single checkpoint to perform pairwise, one-to-many, many-to-one, and unconditional generation while retaining assay-specific structure. Across TEA-seq and DOGMA-seq, scMOD achieves the highest average Pearson correlation among the evaluated multimodal VAE baselines, leads specialized translators in several directions, and improves unconditional RNA and ADT generation over prior-sampled baselines. scMOD provides a unified generative framework for completing and generating multimodal virtual cells.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.