Boundary Signals Bias Reference Competition in Protein Language Models
Abstract
Protein language models can predict a masked amino acid by copying from a similar segment in the same sequence. When two matching segments suggest different residues, what determines their relative support? We study this choice in the 600-million-parameter ESM Cambrian (ESM-C) model (ESMC600M) by changing the boundary position encoding used by one attention head. We block attention to exterior sequence while keeping amino acids and physical positions fixed. Interventions identify a contribution from boundary values that affects later reference keys. We use this route to predict score changes under new boundary interventions. On unseen sequence sources and layouts, the predictor reduces mean absolute error from 0.115 with static input features to 0.093. Removing the boundary contribution attenuates the effect more than norm-matched random removal and preserves copying when the references agree. Natural-input tests show that reduced boundary sensitivity can coexist with worse residue prediction. Boundary stability therefore needs to be evaluated alongside prediction accuracy.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.