PODA: Privileged On-Policy Distillation for Antibody and Nanobody Design
Abstract
Epitope-conditioned de novo antibody design seeks to generate CDR sequences and the corresponding antibody–antigen complex structure for a specified target site. Flow-matching models learn from interpolations toward native complexes, but generate through their own predicted states, creating a potential mismatch between training and inference. We introduce PODA (Privileged On-Policy Distillation for Antibody and Nanobody Design), a two-stage framework for joint CDR sequence and complex-structure generation. Given an antigen sequence and structure, a specified epitope, and a framework sequence with masked CDRs, PODA first trains a deployable student and a privileged teacher with supervised objectives on the same complexes. It then freezes the teacher and updates the student on states obtained by re-noising endpoints generated online by the evolving student policy. Distillation matches teacher-predicted clean endpoints, with residue-level weighting emphasizing CDR and interface geometry and a quality gate selecting teacher targets according to their reconstruction error. On the sequence-filtered test set, PODA increases the design success rate from 20.4% to 32.3% on antibody design and from 12.0% to 24.0% on nanobody design relative to the supervised base model. Ablations support the contribution of privileged supervision, and PODA achieves the highest mean DockQ among the evaluated post-training methods in both cohorts. These findings support privileged endpoint supervision on student-induced states as an effective approach to improving complex interface geometry.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.