Opt-Judger: Fine-Grained Evaluation of LLM-Based Optimization Modeling via Feasibility Verification
Abstract
Large language models are increasingly used to generate optimization models from natural language descriptions, motivating the development of methods to evaluate their correctness. However, existing evaluators still struggle to detect constraint errors in generated models and provide limited feedback for diagnosing them. These limitations hinder accurate assessment of LLMs’ ability to translate problem requirements into optimization models and offer little guidance for improvement. To address these limitations, we propose Opt-Judger, a solver-backed framework for fine-grained evaluation of LLM-based optimization modeling. Opt-Judger evaluates whether the constraints in generated models faithfully capture the problem requirements by checking whether they permit the same decisions as a reference model representing these requirements. It uses feasibility verification to search for a witness point, which represents a decision that is feasible under one model but infeasible under the other. This point reveals constraint violations that indicate mismatches with the problem requirements and supports error localization. We also construct Opt-Judger-Bench using equivalent reformulations and transformations that simulate common modeling errors. Extensive experiments show that Opt-Judger accurately identifies genuine constraint errors across benchmarks and consistently outperforms existing evaluators on challenging cases.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.