acceptodds
Under review as a conference paper at ICLR 2027

Hallucinations in LLM-Based Optimization Modeling: Definition, Taxonomy, Detection Benchmarks, and Reference Detectors

Abstract

Large language models (LLMs) are increasingly used to translate natural-language optimization problems into mathematical formulations and solver code. Most existing work evaluates the resulting models by comparing the objective value produced by generated solver code with a reference value. However, matching the reference objective value is not a reliable test of correctness, as it can mask unintended changes to optimization semantics, such as changes to the intended decisions, feasible region, objective meaning, or solver behavior. We define an optimization-modeling hallucination as a deviation from the intended optimization semantics. We develop, to our knowledge, the first fine-grained hallucination taxonomy specifically for optimization modeling, spanning objective, variable, constraint, and implementation failures. Building on this definition and taxonomy, we formulate optimization-modeling hallucination detection as structured prediction over tuples comprising a problem description, symbolic model, and solver code. To support detector development and evaluation, we construct development and test suites comprising more than 12,000 detection instances, with disjoint source-problem sets. The test suite contains correct examples for measuring false alarms, controlled changes for measuring localization, and LLM-generated outputs for measuring detection of naturally occurring errors. Using these suites, we design and evaluate six detectors for optimization-modeling hallucinations. Among them, OptArgus is a multi-agent detector in which separate agents inspect the objective, variables, constraints, and solver code before combining their findings into a final diagnosis. Experiments show that OptArgus offers a strong overall trade-off across the three evaluation settings, combining the highest controlled-error localization scores among the six detectors with competitive false-alarm control and natural-error detection, while producing concise diagnoses. Together, these contributions establish optimization-modeling hallucination detection as a concrete empirical problem and foundation for more reliable optimization modeling.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.