REASON, EXPLORE, VERIFY: a neuro-symbolic architecture for constrained long-horizon planning
Abstract
Large Language Models (LLMs) can interpret natural-language goals and generate plausible action sequences, but remain unreliable for constrained long-horizon planning. Errors can accumulate across decisions, locally plausible actions can eliminate feasible continuations, and generated plans can violate numerical, logical, or temporal requirements. We introduce Reason–Explore–Verify (REV), a neuro-symbolic architecture that combines hierarchical problem formulation, explicit trajectory exploration, and formal verification of candidate decisions. We instantiate REV as Hierarchical Monte Carlo Tree Search with SMT Gating (HMCTS-SMT), which integrates hierarchical LLM agents, LLM-guided Monte Carlo tree search, and Satisfiability Modulo Theories solving. Each search node maintains an agentic state and a branch-specific symbolic constraint state. A selected action creates a successor only if its symbolic update preserves satisfiability, making symbolic verification a mandatory transition condition rather than an optional tool or post hoc check. We evaluate HMCTS-SMT on supply-chain optimization and travel planning, covering optimization-oriented and feasibility-oriented settings. Across these tasks, HMCTS-SMT improves plan validity and solution quality over direct generation approach and search-based LLM baselines. These results demonstrate the effectiveness of combining hierarchical decomposition, explicit exploration, and transition-level symbolic verification for constrained long-horizon planning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.