acceptodds
Under review as a conference paper at ICLR 2027

Hyper-Fold: Exploring the Expressive Limit of Sequence-Geometry Learning for Proteins via Hypergraph Modeling

Abstract

Protein structure modeling rests on a single computational primitive: the interaction between what a residue *is* (sequence content) and *where it sits* (three‑dimensional geometry). What is the expressive limit of this layer class? We show that the complete bilinear operator over content‑geometry outer products—the sufficient statistic of all second‑order interactions—is the expressive ceiling, while the additive message passing of mainstream geometric GNNs is provably blind to content‑geometry binding. We then introduce **Hyper‑Fold**, a rank‑ separable convolutional backbone approaching this ceiling at message‑passing cost: each radius neighborhood is organized into a sequence hyperedge and a contact hyperedge, modulated by an edge‑conditioned matrix‑valued operator factorized into learned basis operators with geometry‑generated coefficients. Across enzyme function prediction, fold classification, and ligand binding site detection, **Hyper‑Fold** and its hierarchical variant **Hyper‑Fold‑Deep** achieve the best results among protein‑specific structure encoders; **Hyper‑Fold‑Pocket**, an anchored set‑prediction head, surpasses UniSite‑3D on UniSite‑DS and two zero‑shot benchmarks with no sequence language model features, fewer parameters, and lower latency—suggesting that a sufficiently expressive 3D backbone recovers information that fusion architectures previously borrowed from evolution‑scale pretraining.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.