iGC: Antibody Optimization via Reinforcement Learning over a Learned Somatic Hypermutation Prior
Abstract
Antibody affinity maturation is a natural optimization loop. Individual B cells propose mutations through somatic hypermutation (SHM), variants with higher affinity for antigen are positively selected, and the process repeats cyclically, giving rise to a population of high-affinity antibody producing B cells. A computational model of this process should therefore capture both how mutations arise and how they are selected for antigen-specific affinity. However, modeling these two components jointly is challenging: mutation models capture the sequence patterns of SHM but do not directly optimize antigen binding, while affinity optimization can improve binding without accounting for biologically plausible mutation trajectories. We introduce iGC (in silico Germinal Centers), a two-stage reinforcement learning (RL) framework for antigen-specific antibody optimization. Stage1 learns an SHM mutation prior from parent–child pairs in B-cell lineage trees and domain-adapts it to the target affinity-maturation dataset. Stage2 initializes an RL policy from this adapted prior and optimizes antigen binding while remaining Kullback–Leibler (KL) regularized toward a frozen copy of the prior. Direct optimization of affinity rewards can lead to degenerate solutions, including excessive use of a small number of highly rewarding substitutions or convergence toward structurally tolerated framework mutations. To mitigate these effects, we introduce a Reward–Tax Meta-Gradient objective. RewardNet smooths the affinity landscape, while TaxNet discourages reward overexploitation; their difference defines a learned meta-reward for policy optimization. We evaluate whether iGC achieves both objectives: preserving SHM-consistent mutation behavior and improving antigen-specific binding. The mutation prior more than doubles the strongest baseline on mutation-site prediction. Following RL optimization, iGC preserves strong SHM sequence structure, with 57.0% of predicted mutation positions falling within WRC hotspots (canonical DNA motifs preferentially targeted during SHM). iGC achieves a strong exploration–exploitation balance, yielding competitive reward–mutation-diversity products across candidate-selection cutoffs. These results demonstrate that iGC can jointly model biologically plausible mutation generation and antigen-specific affinity selection, providing an in silico framework for iterative antibody affinity maturation. Code and implementation details are available at https://anonymous.4open.science/r/iGC-0A08/README.md.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.