How Much Do Antibody Benchmarks Depend on Related Training Examples? A Cross-Benchmark Audit
Abstract
Machine-learning antibody benchmarks can retain training examples related to antibodies in the test set, obscuring how much reported performance depends on that support. We audit four predictive benchmarks while keeping each test set fixed. Ordinary training O uses the original training pool, family-filtered training F removes examples related to fixed test antibodies under a benchmark-specific rule, and size-matched blind control S removes the same quantity without using family relationships, labels, or outcomes. In AVIDa, O–F is positive in all 160 robustness analyses, and median O–S is approximately zero despite no exact mature-VHH recurrence. In AbDesign, scaffold-averaged O–F is 0.142 with a 95% bootstrap interval of [0.038, 0.259]. AlphaSeq shows large O–F values for AbLang2 and ESM-2, while MutAb shows smaller and more variable differences. The median auxiliary distribution-matched control remains closer to O than F in all four predictive benchmarks. In AbDesign, the same direction persists under IMGT CDRH3/CDRL3 and global-alignment relations fixed before evaluation that alter realized family memberships. Between ESM-2 and AbLang2, the preferred model remains unchanged across O, F, and S in three benchmarks and changes only in AlphaSeq. Antibody benchmarks therefore differ substantially in their dependence on related training support.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.