Multimodal Latent Flow Matching for De Novo Nanobody Design
Abstract
Generative protein design must leverage heterogeneous sequence and structural data while supporting precise control over molecular interactions. We intro- duce Proteina Mixta, a multimodal latent flow-matching model for all-atom co- design and sequence-only generation. A shared autoencoder maps sequence-only, backbone–sequence, and all-atom inputs into continuous residue-level represen- tations, enabling one model to generate across modalities. By modeling dis- crete amino acid sequences through continuous latent dynamics, Proteina Mixta achieves strong performance in both all-atom monomer co-design and sequence- only generation. We extend Proteina Mixta to target-conditioned nanobody (a single-domain antibody) design by assigning synthetic annotations of antibody framework and complementarity-determining region (CDR) to general protein interfaces. Antibody-specific conditioning is then learned entirely from these synthetic examples: true CDR and framework labels are never seen during rep- resentation learning or generative training. The multimodal formulation further enables direct co-training of the structure-based generator with sequence-only antibody repertoires from the Observed Antibody Space (OAS) database. Training on liability-filtered OAS sequences increases effective CDR3 diversity by 26% and reduces sequence-liability rates by 13%. In a controlled comparison with common targets, VHH frameworks, CDR lengths, and sampling budgets, Proteina Mixta achieves up to 50-fold higher unique in silico success rates than prior open models, without sequence redesign. These results suggest that a shared latent space lets structure-free data supervise all-atom generation, opening a route to de novo structure design at the scale of sequence data.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.