acceptodds
Under review as a conference paper at ICLR 2027

Can Post-Training on Agentic Chain-of-Thought Autonomously Screen Protein Binders For Experimental Success?

Abstract

A well-known bottleneck exists in protein binder design: favorable in silico metrics do not necessarily translate to guaranteed success in the wet-lab. To address this, we introduce BinderReason: we post-train large language models (LLMs) via a two-stage post-training approach on Chain-of-Thought (CoT) reasoning traces generated by our autonomous multi-agent system, BinderLab-Think. We compile a curated collection of 6650 design-target complexes from three open-source datasets, representing 40 targets, of which 18.40% of designs are experimentally validated as good binders. These complexes are co-folded and analysed by BinderLab-Think, our closed-loop system of 11 specialised agents using Plan-and-Solve CoT prompting to predict whether each design binds to its target. We cold-start two open-weight models on the resulting traces to understand the biological, chemical, and physical concepts involved in binder analysis, and to align their internal thinking with a critical, systematic, and motivated approach to verdict generation. We then refine the cold-started Qwen3.5-9B model, BinderReason-SFT, via Group Relative Policy Optimisation (GRPO) with a domain-specific composite reward comprising five components: metric-weighted correctness, reasoning consistency, non-degeneracy, label validity, and format. This reward encourages consistency with structural evidence, suppresses hallucinated metric values and degenerate outputs, and enforces a parsable verdict. The resulting model, BinderReason-RL, achieves 0.8710 balanced accuracy and 0.8517 F1 on a target-disjoint test split, exceeding the strongest of ten zero-shot closed- and open-weight reasoning baselines by 31.2 percentage points in balanced accuracy. Finally, we analyse how post-training affects both classification and generated reasoning. Although performance remains target-dependent, BinderReason represents a step towards interpretable, domain-informed prediction of experimental binder success.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.