CROW: Constraint-Guided Agentic Workflows for Mathematical Reasoning
Abstract
Large language models have demonstrated strong performance on mathematical reasoning tasks, yet reasoning trajectories that violate critical conditions required for a valid solution can be difficult to correct. Existing approaches primarily focus on improving reasoning generation, but provide limited mechanisms for enforcing the constraints that should govern the reasoning process. We argue that critical reasoning constraints provide explicit conditions for guiding and verifying reasoning, particularly on challenging mathematical problems. We introduce Constraint Reuse and Orchestration Workflow (CROW), a constraint-guided agentic workflow for mathematical reasoning. CROW formulates constraint utilization as a sequential decision-making problem, allowing the orchestrator to selectively retrieve, inspect, reuse, disregard, or propose constraints according to the current problem. The orchestration policy is optimized with reinforcement learning based on downstream solution correctness. During training, newly proposed constraints that contribute to successful solutions are retained for future use, progressively expanding the set of available constraints. Through reinforcement learning (RL), the orchestrator learns when to acquire and apply the retrieved constraint, and when to disregard it due to limited relevance or applicability to the current problem. Experiments on HARDMath, HARDMath2, and DeepMath show that CROW substantially improves performance on HARDMath, where critical constraints contribute to the largest gains, while RL-trained orchestration preserves strong performance across the broader benchmark suite.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.