AstroMathBench: Evaluating the Mathematical Reasoning Capabilities of LLMs for Astronomical Problems
Abstract
Large language models (LLMs) have demonstrated remarkable capabilities in advanced mathematical reasoning tasks. However, computational reasoning in astronomy with LLMs presents unique challenges yet has received much less attention. Existing benchmarks often fail to capture LLMs' cross-disciplinary reasoning abilities and are prone to data contamination. To address these problems, we propose AstroMathBench, a benchmark for evaluating LLMs' computational reasoning capabilities in astronomy. AstroMathBench covers five task types in astronomical reasoning: formula calculation, formula derivation, unimodal computation, multimodal computation, and code-based calculation. Beyond task-level evaluation, we decompose the reasoning process into five cognitive dimensions, enabling fine-grained assessment of LLMs' capability boundaries. Moreover, we design a cross-validation-based pipeline to mitigate data contamination by separating private and public data and employing isomorphic variant generation. The resulting dataset contains 554 seed problems that have been manually verified and checked for contamination, with 90.1% fully code-formalized. Our evaluation of 10 leading LLMs finds that the highest accuracy reaches only 76%, highlighting the significant challenges of computational reasoning in astronomy and the need for LLMs with stronger cross-disciplinary reasoning abilities. In general, we hope that AstroMathBench will drive advances in AI for astronomy.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.