acceptodds
Under review as a conference paper at ICLR 2027

Guiding protein ensemble generators with cryo-EM particles, without maps or poses

Abstract

Understanding a protein's function requires knowing its conformational distribution: what forms it takes and how often. Cryo-electron microscopy provides this distribution in principle, since every particle images a single molecule, but the averaging needed to address the experiment's low signal-to-noise ratio throws away the very heterogeneity we seek. Pretrained generative models of protein ensembles could restore part of what is lost if individual particles steered their sampling. However, this steering would require a likelihood of the observed images given a generated structure, and constructing this term requires knowledge of the particle's pose. We build this likelihood in a low-dimensional code shared by a frozen pretrained generative model for protein conformations and a frozen cryo-EM image encoder, and fit the code with simulated structure–image pairs. Because the code holds only information available in both models' representations, it suppresses imaging parameters such as defocus. The scatter that imaging, pose included, still leaves between images of the same conformation sets the likelihood's covariance. Using this likelihood, we guide generation with only 128 particles and no density map or pose estimates toward rare states and PDB structures. This is possible because the prior fills in high-resolution details, while the particles constrain large-scale rearrangements. The same code also recovers free-energy landscapes by deconvolving imaging scatter, resolving barriers that particle counting cannot. We test the approach on simulated particles of a GroEL subunit, lysenin and DppA.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.