acceptodds
Under review as a conference paper at ICLR 2027

DARPinBench: A Dataset and Library-Screening Benchmark for Designed Ankyrin Repeat Proteins

Abstract

Pulling a handful of experimentally confirmed binders to the front of a library whose members differ by one or two point mutations is a recurring machine-learning problem, from antibody affinity maturation to the triage step of a generative design pipeline. Designed ankyrin repeat proteins (DARPins) instantiate it: one shared solenoid scaffold, sparse positives, and negatives that are near neighbors. Public sources keep sequences, complex coordinates, and binding readouts apart, so neither an auditable table nor a frozen ranking protocol exists for this scaffold. We release both. DARPinDB annotates sequence, binding label, , and evidence source on a single record: 188 records, 179 binders, and 37 targets, of which 121 carry a numeric , with structures tiered and released with chain maps. DARPinBench freezes three screening modes on that table. Mode 1 scores ten mainstream methods from the sequence, structure, rescoring, and docking families on ten near-neighbor libraries of 500 sequences each; Modes 2 and 3 release a shared 20,000-member library and a leave-one-target-out split as held-out material, unscored here. No method reaches the top 1% of a library reliably: the best rate at the top 1% is of ten units, AF3 attains the highest recall at the top 10% () while reaching only at the top 1%, and the median enrichment factor at 5% is zero for eight of the ten methods. A high DockQ does not travel with a front-of-library rank. We introduce no new scorer; the contribution is a frozen reference line, together with a diagnosis of where off-the-shelf scores fail, for later evolution and training to exceed.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.