acceptodds
Under review as a conference paper at ICLR 2027

Multimodal Latent Flow Matching for De Novo Nanobody Design

Abstract

Generative protein design must leverage heterogeneous sequence and structural data while supporting precise control over molecular interactions. We intro- duce Proteina Mixta, a multimodal latent flow-matching model for all-atom co- design and sequence-only generation. A shared autoencoder maps sequence-only, backbone–sequence, and all-atom inputs into continuous residue-level represen- tations, enabling one model to generate across modalities. By modeling dis- crete amino acid sequences through continuous latent dynamics, Proteina Mixta achieves strong performance in both all-atom monomer co-design and sequence- only generation. We extend Proteina Mixta to target-conditioned nanobody (a single-domain antibody) design by assigning synthetic annotations of antibody framework and complementarity-determining region (CDR) to general protein interfaces. Antibody-specific conditioning is then learned entirely from these synthetic examples: true CDR and framework labels are never seen during rep- resentation learning or generative training. The multimodal formulation further enables direct co-training of the structure-based generator with sequence-only antibody repertoires from the Observed Antibody Space (OAS) database. Training on liability-filtered OAS sequences increases effective CDR3 diversity by 26% and reduces sequence-liability rates by 13%. In a controlled comparison with common targets, VHH frameworks, CDR lengths, and sampling budgets, Proteina Mixta achieves up to 50-fold higher unique in silico success rates than prior open models, without sequence redesign. These results suggest that a shared latent space lets structure-free data supervise all-atom generation, opening a route to de novo structure design at the scale of sequence data.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.