From Demonstrations to Solver-Guided Exploration: Learning to Formulate Constraint Programs from Natural Language
Abstract
Translating natural-language descriptions into executable constraint programming (CP) models requires more than generating code that compiles. A model may return a solution while omitting a required constraint, using an incorrect domain, or optimizing the wrong objective. Human experts address such errors by inspecting solver feedback and revising the formulation, whereas existing LLM-based CP modeling systems often generate a model once or follow a fixed repair procedure. We introduce CPARL, a solver-grounded learning framework that combines CP-specific supervised fine-tuning with agentic reinforcement learning. The supervised stage first teaches direct modeling and solver-assisted revision through general demonstrations, followed by a second round targeting errors identified on a separate development set. Building on these capabilities, the reinforcement learning stage learns to choose between direct generation and solver-assisted revision, and to control model edits, solver operations, and termination. For tool-assisted trajectories, we propose Solver-Event Group Relative Policy Optimization (SE-GRPO), which groups decisions between successive solver evaluations and assigns credit using incremental solver evidence, anchored by semantic validation. We further introduce NL4CP, a new CP modeling benchmark built from formally specified problems, and evaluate CPARL on NL4CP and four public benchmarks. Using a Qwen3.5-27B backbone, CPARL achieves 35.0% micro and 44.6% macro pass@1 accuracy, exceeding all evaluated baselines fine-tuned for optimization modeling and directly prompted DeepSeek-V4-Pro and GLM-5.1 in both micro and macro averages. Ablation studies examine the contributions of the training stages, supervision mixture, and solver-event credit assignment. We analyze code executability and executable-but-invalid failures.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.