acceptodds
Under review as a conference paper at ICLR 2027

Protein Localization & Identification Directly from CryoEM Maps

Abstract

Cryo-electron microscopy (cryoEM) produces three-dimensional maps of the electrostatic potential of protein complexes. Interpreting a map requires building an atomic model, which in turn requires knowing which protein occupies each region. Automated methods answer this by tracing the protein backbone and identifying side chains, and both steps fail below 5Å resolution. In this work, we propose a method that assigns proteins to regions of the map without having to build an atomic model. We align a frozen foundation model of cryoEM maps with a frozen protein language model in a shared embedding space, training only two lightweight heads on cached features. Two components are essential: a soft-target contrastive objective that treats solvent as an explicit class and distributes target mass over every chain of a complex, and a rigid-motion-invariant attention module over the sampled voxels of a map. Given a map and the sequences of the proteins it contains, our model assigns regions of the map to their corresponding proteins, outperforming controls. This performance is retained at resolutions where backbone tracing is no longer reliable. When the protein composition is unknown, the same embedding can instead retrieve candidate sequences from a large catalogue, placing the correct protein among the top-ranked candidates in more than 84% of cases. Together, these results show that protein identity can be inferred directly from cryoEM density without first constructing an atomic model.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.