acceptodds
Under review as a conference paper at ICLR 2027

VASIC: Validation-Guided Adaptation for Symbolic Integration with LLMs and CAS

Abstract

AI systems are increasingly used in research and education, generating complex proofs, derivations, and results faster than experts can reliably validate. Interactive theorem provers such as Lean provide strong guarantees for formalized proofs, but fully formalizing the symbolic derivations and analytical calculations central to physics and engineering is often impractical. As a result, a growing volume of LLM-generated mathematics in these domains remains unverified. To bridge this gap, we propose a complementary, scalable method based on Computer Algebra Systems (CAS) to improve the reliability of LLM-generated mathematical outputs. As a proof of concept, we present VASIC, a validation-guided LLM-CAS protocol for symbolic integration. VASIC tests each candidate LLM solution using symbolic and numerical methods. Failed attempts are routed through feedback-guided adaptation and retried. We benchmark VASIC against standalone CAS and LLM baselines from a broad range of model families and versions. Compared with direct use of the same model, VASIC increases validated success rates by 3.5–10 percentage points for frontier models and by up to 36.2 percentage points for smaller and less expensive models. For example, DeepSeek Reasoner V3 achieves 84.0% validated success with full VASIC, close to DeepSeek V4's 85.2%, at roughly half the estimated model-inference cost, counting the additional LLM requests made by VASIC. Together, these results show that LLM-CAS hybrids can increase AI success rate and practical reliability in problems of general importance for science and engineering.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.