acceptodds
Under review as a conference paper at ICLR 2027

Does Reasoning Improve LLMs Beyond Mathematics? A Biological Test on Intrinsically Disordered Protein Regions

Abstract

Reasoning improves language models in domains such as mathematics, where premises and answers are well defined. Whether these benefits extend to biology, where evidence is incomplete and heterogeneous, remains unclear. We study this question through prediction of intrinsically disordered protein regions (IDRs). GPT-5.4 generated evidence-grounded teacher traces that were transferred to Qwen3-4B through supervised fine-tuning (SFT) and Group Relative Policy Optimization (GRPO). We combined the resulting models with a frozen ESM-2 residue classifier and calibrated interval decoder. Zero-shot Qwen produced valid structured outputs but no usable IDR intervals, showing that formatting alone does not confer biological prediction ability. Both hybrid models performed strongly on teacher-corpus validation and retained substantial performance on an independently selected, sequence-disjoint test set, numerically exceeding two leading CAID2 Disorder-NOX predictors evaluated with the same labels. GRPO improved interval localization, trace–answer consistency, and reasoning coherence, but did not consistently improve residue classification over SFT. Perturbation analysis identified candidate-region evidence as the dominant predictive signal, while claim-level auditing showed that coherent traces could still contain unsupported statements. Overall, our reasoning-based framework improves selected biological prediction objectives under incomplete evidence, but its benefits are objective-dependent; the strongest results arise from combining language-model reasoning with specialized protein representations and explicit grounding checks.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.