acceptodds
Under review as a conference paper at ICLR 2027

Distillation, Self-Training, and the Fragility of Learned Chemical Reasoning

Abstract

Building specialized models for scientific reasoning presents a fundamental tradeoff: frontier models provide strong domain capabilities, but continued reliance on them for supervision limits the accessibility and adaptability of smaller open-weight models. We investigate whether frontier supervision can instead serve as a one-time bootstrap, using forward reaction prediction, a core task in computer-aided synthesis planning, as a testbed. We introduce Resonance, a post-trained model that distills frontier-model reasoning into a smaller student model and then removes the teacher while continuing to train on the student's own verified and preference-ranked generations. Resonance transforms a base model with near-zero accuracy into an effective reaction predictor. Through reasoning perturbations and counterfactual prompting analyses, we provide additional insights into the chemical reasoning acquired by trained student models. Together, our results demonstrate the effectiveness of our pipeline in instilling reaction prediction competencies while also uncovering systematic pathologies in our model's ability to reason over chemical contexts.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.