acceptodds
Under review as a conference paper at ICLR 2027

Unifying Protein Domain Parsing and Classification at Scale

Abstract

Protein domains are modular units of protein structure that play central roles in protein folding, function, and evolution. Annotating a protein structure requires answering two coupled questions: which residues form each domain, and what fold does each domain adopt? Existing methods largely address these questions separately, with domain parsing and fold classification performed by sequential stages. Such pipelines cannot optimize the two tasks jointly and may propagate parsing errors into subsequent classification. We introduce joint protein domain parsing and fold classification and construct TEDBench-Joint, a large-scale benchmark derived from The Encyclopedia of Domains (TED). We further propose DoTR, an end-to-end SE(3)-invariant model that encodes residue-level 3D structure and uses a transformer decoder to directly predict an unordered set of domains, each represented by a residue mask and a fold label. Experiments on TEDBench-Joint show that DoTR outperforms dedicated domain-parsing methods on fold-agnostic parsing metrics and two-stage parsing-and-classification pipelines on fold-aware metrics. These results demonstrate that domain parsing and fold classification can be effectively learned within a single model, providing a unified approach to protein domain annotation at scale.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.