Humans vs. Autonomous Agents in De Novo Protein Binder Design: A Post-Competition Analysis of the muni x Adaptyv TREM2 Hackathon
Abstract
We present, to our knowledge, the first experimentally validated comparison of human-led and autonomous-agent de novo protein binder design. In a single-day hackathon, human teams and six agents submitted 81 and 60 designs under common constraints; 100 designs from nine human teams and six agents were tested. Of these, 89 expressed and 37 displayed binding to TREM2. Human and agent hit rates were 25/65 (38.5%) and 12/35 (34.3%), respectively (two-sided Fisher’s exact P=0.83), with no clear cohort advantage. The strongest binders had per-design mean values of 1.113 nM (human) and 3.636 nM (agent) by SPR. Experimental stability measurements were not available for the tested designs. Sequence-validated Protenix and Chai-1 predictions placed the strongest binder and all 36 mini-protein hits at TREM2’s hydrophobic tip, outside the principal crystallographic contact sites of scFv-2 and scFv-4; 46/49 expressed mini-protein non-binders were also predicted to contact this region. Retrospective Protenix ipSAE and ipTM distinguished binders from non-binders (AUROC 0.782 and 0.791) but correlated weakly with affinity among 36 fitted binders (Spearman rho=0.293 and 0.196), supporting binder triage but not reliable affinity ranking. A separate 15-round, human-directed agentic search produced ten candidates; nine yielded fitted BLI mean values below 10 nM and three below 1 nM. In this setting, the follow-up produced numerically stronger reported top-end affinities than the original hackathon, although the campaigns differed in duration, selection and assay. Anthropic’s subsequent Claude Science study selected TREM2 as a campaign target and benchmarked its autonomous binder-design results against the publicly released hackathon and agentic follow-up datasets. Together, these results establish the first experimentally validated benchmark of human and agentic protein design, showing similar observed hit rates under common submission constraints, while also highlighting the value of iterative multi-round agentic search.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.