Beyond Reconstruction: Likelihood-Free Reinforcement Learning for Virtual Cells
Abstract
Virtual cell models predict how cells respond to genetic perturbations. A central use of these predictions is deciding which perturbations can shift disease-relevant pathways in a desired direction and accelerate therapeutic target discovery. Because these models are trained to reconstruct gene-level expression, they can reduce reconstruction error without ranking the perturbations with the strongest pathway effects near the top. Pathway-level effects are defined over gene sets, relative to a background set, rather than individual genes so they are difficult to express as supervised losses but can be directly scored as rewards, which suggests benefits from reinforcement learning (RL). Standard RL, however, requires policy likelihoods, whereas virtual cell models generate high-dimensional cell populations without tractable likelihoods. We introduce Maximum Mean Discrepancy Policy Optimization (MMD-PO), a likelihood-free reinforcement learning method that weights rollouts sampled from a frozen policy by biological rewards, and moves the current policy toward this reward-weighted empirical target by minimizing maximum mean discrepancy between rollouts. We pair MMD-PO with Gene-Effect and Pathway-Enrichment rewards. In a controlled synthetic benchmark with known latent effects, MMD-PO improves held-out ranking over continued supervised training, GRPO, and online DPO, and its gains depend on feedback reliability, candidate diversity, and the headroom of the starting model. On held-out perturbations in HepG2 cells and zero-shot transfer from resting to stimulated primary T-Cells, MMD-PO improves top-% hit enrichment over GRPO by % and %, respectively, while retaining prediction fidelity. These results establish likelihood-free policy optimization as a practical interface between virtual cell models and experimental prioritization.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.