acceptodds
Under review as a conference paper at ICLR 2027

Peppermint: structure-conditioned representation learning for interacting protein pairs

Abstract

Protein-protein interactions depend on the sequences of both proteins and their spatial arrangement, motivating protein representations that incorporate molecular context. We introduce Peppermint, a structure-conditioned masked language model for interacting protein pairs. Peppermint maintains separate chain representations and alternates intra-chain self-attention with bidirectional cross-attention, preserving chain identity while allowing each residue to incorporate information from the other chain. The architecture assigns distinct roles to local conformation and interaction geometry: gated embeddings of discrete structural states and continuous backbone geometry enrich residue representations, while directional geometric biases guide attention between chains. Initialized from a pretrained sequence encoder, the model learns from paired sequences and protein complexes through masked amino-acid modeling and auxiliary structural prediction. Mixing examples with and without structure allows the same model to operate under either input condition. Supplying structure improves simultaneous recovery of masked interface residues by 9.3 percentage points relative to sequence-only input, with downstream benefits varying by task and metric. These results support incorporating complex geometry into paired protein representations while retaining sequence-only inference within the same encoder.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.