TreasuryOpt: Benchmarking LLMs on Generating Reliable Optimization Programs
Abstract
Large language models (LLMs) are increasingly explored for mathematical optimization, yet existing evaluations rarely test whether they can construct reliable optimization programs that remain robust across unseen instances. We introduce TreasuryOpt, a benchmark evaluating LLMs on generating reliable optimization programs grounded in enterprise treasury operations, comprising four scenarios, 28 modeling cases, and 596 hidden test instances. For each modeling case, a system must construct a fixed program from its textual requirements, which is then evaluated across unseen instances without further model intervention. Our solver-verified evaluation protocol then extracts the returned decision variables and independently recalculates constraint satisfaction and objective optimality against solver-certified reference bounds. Extensive experiments across six frontier LLMs—under standalone and agentic settings across solver-hidden and solver-aware conditions—reveal a persistent feasibility–optimality gap: systems face considerable difficulty satisfying domain constraints, with success rates collapsing further when requiring solver-certified optimality. Furthermore, while interactive coding harnesses enhance solution quality via iterative self-debugging, they incur orders of magnitude higher generation costs with only limited performance gains. These results demonstrate that reliable optimization with LLMs requires progressing beyond surface-level executability toward mathematically sound and generalizable program generation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.