acceptodds
Under review as a conference paper at ICLR 2027

Does Register-Length Flexibility Help Peptide–MHC-II Classification? A Controlled Ablation on the Chicken Allele BLB2

Abstract

Discrete latent-alignment models are usually given a fixed-length candidate structure to select over. We ask what happens when that support is widened to match documented heterogeneity in the target structure itself, using peptide–MHC class II (MHC-II) ligand classification as a test bed. Most MHC-II predictors represent the binding core as a canonical nine-residue window and treat only its position as a latent register variable. The chicken allele Gaga-BLB2*002:01 is a useful test case because a structural study establishes that it can anchor a decamer core (Halabi et al., 2021), while a separate computational motif-deconvolution study reports its ligand pool splits nearly evenly between 9-residue and 10-residue annotated cores (Racle et al., 2023), a split our own compiled dataset reproduces (2,134 9-mers, 2,050 10-mers among 4,184 positives). We train a register variable, supervised by an auxiliary cross-entropy loss against the annotated core alongside the primary classification loss, to range jointly over both candidate lengths, and compare it under a parameter-matched ablation that masks the 10-mer candidates out. Holding architecture, data, and training procedure fixed across five outer-fold held-out evaluations (single seed, one allele), the two variants are statistically indistinguishable on ranking metrics (pooled ROC-AUC 0.920 for both), and the restricted model is, if anything, slightly higher on PR-AUC (0.597 versus 0.583; a paired bootstrap places the gap at 0.013, 95% CI [0.006, 0.021]). We cannot attribute that gap to candidate-length support alone, since masking also removes auxiliary supervision from the 10-mer-annotated half of positives. Only the flexible model can agree with a length-flexible annotation at all: it matches the annotated core length on 74.3% of held-out positives (a 51.0% majority-class floor) and reproduces the exact annotated core substring on 30.7%. We also find that a large share of what our dataset labels "BLB2-positive, 9-mer core" may in fact be a co-expressed allele's ligands misassigned by the motif-deconvolution pipeline, a hypothesis Halabi et al. themselves supply and that predicts the largest effect we observe: 9-mer-annotated positives score far lower (PR-AUC ≈0.34–0.35) than 10-mer-annotated positives (PR-AUC ≈0.58–0.59) at matched prevalence. Register flexibility, in this architecture and on this single allele, is therefore better read as a lever on whether a model's internal explanation can engage with heterogeneous annotations at all, not as a classification-accuracy improvement. Part of the negative result may itself be an artifact of imperfect training labels rather than of the inductive bias being tested.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.