Multi-Objective Amortized CDRH3 Sequence Sampling Diversifies Promising Antibody Candidates
Abstract
Antibody therapeutics have transformed treatment across diverse therapeutic areas and have become central to modern medicine. While recent machine learning approaches aim to accelerate antibody discovery through structure-based design and target-affinity optimization, successful candidates must also satisfy broad developability constraints, including structural compatibility, low immunogenicity, and favorable biophysical properties, while remaining sufficiently diverse to provide sufficient coverage of design space to mitigate later-stage development failures. We introduce a preference-conditioned multi-objective generative flow network (GFN) fine-tuning pipeline for fixed-framework CDRH3 loop redesign that fine-tunes a biological sequence model against many inexpensive proxies covering backbone compatibility, sequence likelihood, immunogenicity risk, novelty, and developability liabilities. It combines up to fourteen predictive rewards in a single policy trained to sample CDRH3 loops across 623 SAbDab antibody structures, leveraging a reward reduction strategy to preserve Pareto dominance under this learned map. We also introduce a novel autoregressive span-in-filling variant of AMPLIFY-350M that serves both as a scaffold-conditioned CDRH3 generation prior and as a direct conditional likelihood scorer. We evaluated reward scores, diversity, and design quality on loops sampled on 45 held-out antibody complexes to assess structural consistency, confidence, aggregation propensity, stability, and antigen binding. Compared with the inverse-folding methods ProteinMPNN and ESM-IF, our approach produces roughly more distinct acceptable candidates per design at 50% sequence identity, while holding of the 14-objective Pareto front pooled across methods, trading per-sample quality for greater exploration of novel possible structures. The results suggest that by relying on a more diverse data distribution and reward distribution, we are able to generate a wider array of promising candidates. All code and data will be made publicly available upon publication.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.