The microscope is the mask: privileged views and labels from a cryo-ET forward model
Abstract
In this work, we use simulated data to train a model for protein annotation in crowded cryo-electron tomography volumes reconstructed from images collected at limited tilt angles and severely corrupted by the measurement operator. Firstly, we leverage the corruptions imposed by the forward model to generate domain-specific augmented views of the exact same scene for an invariance objective integrated into the LeJEPA self-supervised training framework. Secondly, we use additional information from the simulation pipeline such as the positions and identities of proteins in the simulated volumes to inform the architecture of the model and the loss function, so that semantic information is localised at protein positions in the dense feature volume. The resulting model, CARNIVAL, is evaluated without finetuning on classification and detection tasks on a multi-protein benchmark dataset and two in-situ datasets spanning three processing types. We show that CARNIVAL outperforms a state-of-the-art model trained using a contrastive objective on simulated data with a single labelled protein per input block, rather than dense supervision over multiple proteins in the scene, in both tasks on the benchmark dataset, matches or outperforms it in in-situ classification, and computes feature volumes seven times faster. A nine-arm ablation across the five evaluation settings decomposes the mechanism, showing that scene-level supervision is essential and pointing to how the views interact as the source of further gains.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.