acceptodds
Under review as a conference paper at ICLR 2027

LAO-Eval: A Diagnostic Benchmark for LLM Capabilities in Structure-Based Ligand Affinity Optimization

Abstract

Structure-based small-molecule ligand affinity optimization—the iterative editing of hit compounds to improve protein-binding affinity under binding-pocket constraints—is a fundamental task in computer-aided drug discovery. Despite rapid progress in scientific reasoning, the capabilities of large language models (LLMs) across the component steps of this task remain poorly understood. We first show through end-to-end experiments that frontier LLMs achieve only marginal affinity improvements and frequently degrade higher-affinity starting ligands. To identify where these failures arise, we introduce LAO-Eval, a diagnostic benchmark for Ligand Affinity Optimization that decomposes the reasoning process into three independently evaluable stages: Structural Perception, Strategy Discovery, and Strategy Execution, supplemented by chain-of-thought quality analysis. We evaluate 14 frontier LLMs on 10 diagnostic tasks spanning 30 DUD-E protein targets. Our results show that: (i) LLMs perform well on 2D molecular structure recognition (micro-F1 = 0.92) and protein secondary-structure classification (accuracy up to 90%), but their performance drops sharply on 3D ligand–protein interaction perception (F1 0.40); (ii) proposed strategies are usually chemically plausible (93%), and most models generate valid molecules in more than 95% of completed responses, yet at most 25% of intended interactions are realized after Boltz-2 co-folding and structural analysis; (iii) performance differences across model families narrow substantially on interaction realization, indicating a shared limitation of current LLM paradigms; and (iv) controlled distance perturbations show that LLMs respond to spatial information but cannot reliably translate it into interaction predictions or structurally effective edits. These findings identify 3D ligand–pocket reasoning as the principal bottleneck in LLM-based ligand affinity optimization and motivate the development of molecular foundation models with spatially grounded representations and physics-aware structural feedback. Code, data, and reproduction instructions are included in the supplementary reproducibility archive.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.