Does Small Language Models Achieve Scientific Reasoning via Verification-Augmented Reasoning?
Abstract
Small Language Models (SLMs) are a constrained subset of Large Language Models (LLMs) and offer a promising pathway toward deployable models, particularly for scientific reasoning in resource-limited and real-time settings. When integrated with Lean-based formal reasoning, where scientific reasoning is expressed as formal proofs in Lean 4, Coq, and Isabelle, and counterexample-based falsification, where incorrect hypotheses are falsified by admissible counterexamples, their limited capacity raises fundamental gaps in reliable scientific reasoning. Prior works have attempted to address these gaps via Chain-of-Thought (CoT), Self-Consistency, and Tree-of-Thought (ToT), as well as benchmark-driven evaluations including GSM8K and ProofNet, but offer only partial solutions. Recent benchmarks show that even State-of-the-Art (SOTA) SLMs produce incorrect intermediate reasoning steps on challenging problems, while Lean-based formal reasoning and counterexample-based falsification remain underexplored. To address this gap, we benchmark whether verification-augmented reasoning can compensate for limited model capacity. As no suitable dataset exists for this setting, we construct LEANDATA by extending SCIBENCH via the Draft-Sketch-Prove paradigm to translate natural-language scientific hypotheses in 10 domains with Lean-verified proofs and generate admissible counterexamples. We use LEANDATA to train five SLMs (1B-3B parameters) for Lean 4 tactic generation within a neuro-symbolic loop, where the model proposes tactics and the proof assistant provides deterministic execution feedback via reinforcement learning with rulebased rewards. Empirically, +RL+Verification improves few-shot E2E success from 26.8% − 41.9% to 47.3% − 58.9% and the average across VA, FR, SC, AC, and E2E from 33.8% − 46.5% to 51.4% − 61.6%, with VA reaching 48.6%.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.