Towards Symbolic Tool-Integrated Logical Reasoning in Large Language Models
Abstract
Logical reasoning requires deriving valid conclusions from given facts and rules and serves as a foundation for reliable decision-making. Recent advances in large language models (LLMs) have improved their reasoning capabilities, but their reliability under complex logical constraints remains insufficiently evaluated. Existing benchmarks typically involve only a few entities and clues, leaving complex reasoning underexplored. To bridge this gap, we introduce -, a benchmark of 944 expert-verified multiple-choice problems, averaging 8.31 entities, 12.47 attributes, and 18.39 clues per problem. Evaluating 10 representative LLMs in the direct-answer setting shows that complex logical reasoning remains challenging, with the best-performing model, GPT-5.5, achieving only 52.65% accuracy. We further identify Lost in Complexity, a failure mode in which hypothesis enumeration becomes prolonged and unreliable as complexity increases. To address this issue, we propose Tool-Integrated Monte Carlo Tree Search (-MCTS), a post-training framework that trains LLMs to formulate problems symbolically and delegate complex inference to external solvers. By exploring symbolic formulations and solver interactions, -MCTS generates solver-verified trajectories and process-level rewards for autonomous tool use and feedback-driven revision. Based on -MCTS, we develop --8B and --14B, achieving 53.60% and 62.71% accuracy on -, respectively, with --14B outperforming all evaluated methods. Further experiments show that this capability generalizes to other logical reasoning tasks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.