Nonlinearity Leaks Position Through the Mean Pool: A Mechanistic Account of Aggregation in Protein Representations
Abstract
Mean pooling of protein language model (PLM) representations is a linear, permutation-invariant operation that is position-blind; the standard fix is attention pooling, which reweights residues non-uniformly. We start from a puzzle: on antimicrobial peptide activity data, a per-token MLP followed by ordinary mean pooling — still permutation-invariant with respect to token order — matches attention pooling and far exceeds linear mean pooling. We identify the mechanism: when a nonlinearity is applied to embeddings that encode position, it creates token–position cross-terms that survive averaging, so the pooled vector becomes sensitive to which residue sits where. We confirm this causally in a controlled synthetic model: a nonlinear-then-mean representation solves a position-relational task perfectly, but scrambling residue order collapses its accuracy from 1.000 to 0.507 (chance), while linear mean pooling is unaffected in both directions. A scaling analysis shows the effect operates at the 20-symbol amino-acid alphabet ( over chance with an adequately powered probe, recovering 13% of the oracle's signal), diminishing only as the vocabulary grows far past biological scale. On GRAMPA, a multi-pathogen peptide dataset, the same signature appears: a per-token MLP and attention pooling both recover a pathogen-selectivity component (, 10-seed CI for attention) that linear mean pooling () and a hand-crafted physicochemical baseline (, CI includes zero) both miss, and this translates to a multi-seed-confirmed doubling of downstream predictive ( for E. coli, 5-seed, disjoint intervals). We situate the mechanism as a domain-specific instance of the Deep Sets representational result, and are explicit that the causal test is established synthetically; the corresponding real-embedding scrambling experiment is specified as the primary follow-up. The reframing is practical: the operative choice is not mean-vs-attention but linear-vs-nonlinear aggregation of position-carrying embeddings.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.