When Does Few-Shot Transfer Learning Work Across Mutation Assays?
Abstract
Protein mutagenesis can generate large numbers of measurements across many proteins, but characterizing any single new protein-assay pair remains expensive. A typical plate has only a few hundred labeled mutants, too few to train an accurate predictor from scratch. Related measurements offer a way to reduce that burden. A model trained on how mutations shift a property in another assay, or in other proteins, may already encode information useful for the new plate. The question is whether that older data can train a predictor for a new assay from only a small target plate, and whether any gain is genuine transfer rather than protein-identity leakage. We study this with three overlap conditions at a fixed 200-label budget. Where source and target assays measure the same proteins, identity itself may carry the signal, so we compare from-scratch training on the plate to sequential transfer. Where the sources share no exact wild-type sequence with the target, a source checkpoint has no shared identity left to exploit; what may still help is a shared initialization that re-specializes from a few target labels. There we apply first-order model-agnostic meta-learning (FOMAML), with a source protein as the task, meta-trained on solubility and thermostability toward a proteolysis-stability (K50) target, beside a random-initialization control. A third condition holds protein identity out of labeled training and spends the whole budget on one protein. Its plate geometry differs, so we report it beside the matched pair rather than on the same axis. When source and target share the same proteins, sequential transfer clearly beats training from scratch on both ranking and calibration. On the identical plate with sequence-disjoint sources, that gain disappears: every arm, including transfer and FOMAML, ranks near chance despite a large source label pool. On a held-out protein with the budget concentrated on that protein alone, transfer again leads from-scratch at every seed and FOMAML clears its random-initialization control, but both remain far below simply fitting the plate. Spreading the same budget across many held-out identities returns every arm to chance. The pattern is that few-shot transfer here is contingent on shared protein identity among the admitted sources (with a remaining confound that relatedness and provenance also change) and on how the labels are spent, not which reuse method is used.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.