Evaluating Physical Intuition in Large Language Models for Structure-Guided Mutation Design
Abstract
Structure-guided protein design requires interpreting molecular geometry and chemistry to anticipate mutation effects. Existing LLM evaluations emphasize sequence-based prediction or tool-integrated workflows, leaving unclear whether models themselves can reason directly from three-dimensional structures. Here, we systematically evaluate this capacity, termed physical intuition, through a partner-contrastive benchmark that holds the protein and substitution fixed across different binding partners. Models inspect local coordinates with or without Python, without dedicated mutation-effect predictors. Although overall accuracy remains modest, most evaluated LLMs recover more complete partner contrasts when affinity improvements are larger, and successful judgments also occur without Python. Structural interventions and evidence audits support selective dependence on mutation-relevant geometry and chemistry, with auditable observations largely grounded in the supplied coordinates. Case studies further show that models can identify relevant interactions and propose mutant-state hypotheses, but struggle to anticipate structural rearrangements and weigh their energetic consequences. Together, these experiments provide, to our knowledge, the first systematic evidence that general-purpose LLMs can use three-dimensional structural evidence to reason about partner-dependent mutation effects, without relying on specialized predictors to supply those judgments.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.