acceptodds
Under review as a conference paper at ICLR 2027

Shared Motifs, Specific Matches: Learning Protein–Glycan Compatibility

Abstract

Predicting protein–glycan binding for unseen molecules requires recognition patterns that transfer across measured pairs. Recent approaches provide increasingly detailed glycan representations, but models must still learn which local structures matter for each protein from labels assigned to complete molecular pairs. We organize this supervision around recurring motifs: an unseen glycan can share local patterns with measured glycans, while their contribution to binding depends on the protein partner. We introduce **MotifMatch**, which learns one shared parameter vector per motif from measurements on all glycans carrying it. Explicit, linkage-aware motif counts combine these vectors into glycan factors, which sequence-derived protein factors rescale through Hadamard matching. We show that gradient coupling through the shared motif parameters factorizes into motif overlap and the alignment of pair-conditioned responses. Fixed-budget experiments show that measurements on glycans with shared motifs improve prediction for held-out glycans. Controlled comparisons further show that the benefit of Hadamard matching depends on the glycan representation. On GlycanML, MotifMatch achieves the highest mean ranking scores for unseen proteins, unseen glycans, and pairs where both partners are unseen. With protein features cached for all methods, MotifMatch trains and infers faster than the fastest baseline in each stage.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.