LatentFold: Continuous MSA Representations without Discrete Homolog Generation
Abstract
Multiple sequence alignments (MSAs) support accurate protein structure prediction by exposing conservation and coevolutionary signals. Existing MSA generative methods largely instantiate these signals as discrete homologous sequences, even though downstream folding models ultimately consume learned MSA representations. We ask whether useful MSA-derived information must take this discrete form. We propose LatentFold, a fully differentiable framework containing LatentMSA, a sequence-conditioned generator that instead represents structure-relevant MSA signals as a collection of continuous latent states. LatentMSA outputs an MSA-shaped latent tensor rather than discrete homologous sequences, producing distinct probe-conditioned states around a dynamic sequence-specific anchor. It is trained with a permutation-invariant statistical alignment objective that matches selected depth-aggregated moments and residue-pair similarity patterns of natural MSA representations, avoiding explicit correspondence between unordered homolog rows. Experiments on FoldBench and CASP15 show that LatentFold outperforms the evaluated discrete MSA-generation baselines across most concat-MSA-depth comparisons, with the clearest shallow-depth gains on FB_PROTEIN and more metric-dependent improvements on CASP15, while requiring no target-specific homology retrieval at inference time. Its differentiable design allows structural-loss gradients to pass through the frozen folding pipeline and update LatentMSA while the PLM and folding backbone remain frozen. These results support continuous latent MSA generation as a trainable alternative to discrete homolog generation for supplying structure-relevant MSA signals. The code is available at: https://anonymous.4open.science/r/LatentFold-2C68
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.