Changes in Problem Presentation Can Break LLM-Based Optimization Modeling
Abstract
Optimization modeling is the process of translating natural language descriptions of optimization problems into mathematical formulations and solver code. Large language models (LLMs) are increasingly used to automate this process. In practice, users may describe the same problem in different ways. LLM-based modeling systems must handle these different but equivalent problem descriptions correctly, especially in high-stakes applications such as healthcare and energy. Yet we empirically found that LLMs can generate correct mathematical formulations and solver code for an optimization problem, but fail when the same problem is described differently. This raises a central question: how robust is LLM-based optimization modeling across different natural-language descriptions of the same problem? However, existing optimization modeling benchmarks focus primarily on accuracy but provide limited evidence of robustness to such changes in natural-language descriptions. To address this gap, we introduce RobustOpt, an extensible framework with two complementary benchmarks for requirement understanding through question answering (QA) and end-to-end optimization modeling (Coding) to quantify this robustness. Specifically, we systematically construct and validate alternative problem descriptions that preserve the underlying optimization problem, then compare LLM performance across these descriptions. On each benchmark, we evaluate the robustness of 14 LLM-based systems: seven based on general-purpose LLMs and seven on models specifically fine-tuned for optimization modeling. The results show that even fine-tuned models can fail on 34.19% of new descriptions of problems they initially handle correctly; failures reach 67.40% among general-purpose models. Further error analysis reveals mistakes in specifying decision variables, objectives, and constraints. These findings highlight robustness to problem presentation as a key requirement for trustworthy LLM-based optimization modeling, motivating its explicit integration into model evaluation and training.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.