Stroma: A Unified Benchmark for Motif Scaffolding Across Protein Generative Paradigms
Abstract
Motif scaffolding, the design of a protein that preserves a predefined motif while generating the surrounding scaffold, is central to protein engineering. Models that perform this task have advanced rapidly across three generative paradigms: backbone, all-atom and sequence generation. Existing benchmarks were built alongside particular model families, assume incompatible input representations, and cover a narrow range of motifs, so models from different paradigms cannot be compared on equal terms. We introduce Stroma, a unified, extensible benchmark of 100 motif scaffolding tasks, each specified at up to five motif resolutions at once, from residue identity to catalytic side-chain geometry, so that every paradigm receives the same task rather than a translation of it. Nearly a quarter are new, covering protein families no prior benchmark included, among them membrane proteins, nanobodies and nucleic-acid-binding motifs. Where an architecture supports them, we evaluate two conditioning modes (indexed and unindexed) and three motif representations (all-atom, tip-atom and theozyme). Success is defined by explicit thresholds that a single predicted structure must satisfy jointly; we test whether rankings survive re-folding under different structure predictors. We benchmark 12 models (15 checkpoints) spanning all three paradigms. Even though all-atom and backbone models perform similarly on average, only all-atom models can take on the more demanding tasks, such as theozyme scaffolding, which backbone models cannot represent. Under minimal-harness conditioning, protein language models underperform both. Thirteen tasks are solved by no model, pointing to opportunities for further work in motif scaffolding.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.