Conservation-Aware Fine-Tuning: Injecting Physical Laws into LLM Reasoning via Preference Optimization
Abstract
Large language models frequently produce physically implausible reasoning—violating conservation laws of energy or momentum—even when their final numerical answers are correct. We introduce Conservation-Aware Fine-Tuning (CAFT), which treats conservation-law satisfaction as a programmatic preference signal for Direct Preference Optimization (DPO): a rule-based verifier labels each sampled reasoning chain as conservation-respecting or violating, and the resulting preference pairs fine-tune the model to prefer physically valid reasoning. The effect is scale-dependent rather than uniform. At 1.5B, CAFT reduces the Conservation Violation Rate (CVR) by 77.8% (5.62%1.25%, ; scenario-clustered ) and raises accuracy by 16.9 pp (40.31%57.19%, ) on 80 mechanics scenarios; the CVR reduction replicates directionally across three training seeds (1.25%/2.19%/3.75%), while the accuracy magnitude does not. At 8B no arm separates from the base model under our strip-first evaluation protocol, which removes hidden thinking content before verification (all arms at 0/320 violations), so we report the 8B scale as a boundary rather than a transfer. At 32B the base model already sits at the violation floor (0.31%) and CAFT changes nothing measurable ()—a ceiling effect with scale and model family co-varying. These three points distill into a deployment heuristic: apply CAFT only when the base model's CVR confidence interval excludes zero—a rule that is itself only as trustworthy as the protocol measuring CVR. Ablations attribute the conservation benefit to synthetic injection pairs, the accuracy benefit to wrong-answer pairs, and show the benefit is family-local; and a second, independently validated constraint probe shows no measurable side effect.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.