acceptodds
Under review as a conference paper at ICLR 2027

Corrupted Inputs Make Dense Supervision Informative for Zero-Shot Variant Scoring: A Controlled Factorial Study of Compact Protein Encoders

Abstract

Masked language modeling, the default pretraining objective for protein encoders, fixes two separable choices: the input (a sequence with mask tokens, or a complete sequence with some residues replaced by substitutes) and the positions at which the loss is computed (selected or all positions). We cross them with architecture, corpus, optimizer and training length fixed, train compact encoders (7.5M and 33.5M parameters) from scratch on 300k Swiss-Prot sequences, and evaluate zero-shot variant effect prediction on 201 ProteinGym substitution assays. The factors interact: over three seeds at 7.5M, computing the loss at all positions increases mean Spearman by +0.0095 with masked input and by +0.0830 with corrupted input. With masked input almost every visible residue is correct, and the added loss is minimized by copying. Most of the aggregate interaction arises in assays mixing substitution counts, where the count alone predicts fitness; among variants with the same number the interaction is smaller (+0.0194) but positive in every seed, and it remains (+0.0178) when the corruption rate is matched and no pretrained model is used. In the matched design, with loss at all positions, corrupted input outperforms masked input by +0.0890. Under a supervised probe of frozen embeddings, the interaction is not detected in the original design and small in the matched one, so the conclusions concern zero-shot scoring, not learned representations. Sampling substitutes from BLOSUM62 recovers most of the improvement over uniform sampling that a frozen ESM-2 650M teacher gives, and an entropy-preserving permutation of the matrix recovers little.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.