acceptodds
Under review as a conference paper at ICLR 2027

OR-License: Evaluating LLM Agents as Real-World Operations Research Practitioners

Abstract

Can large language model (LLM) agents deliver expert-level operations research (OR) solutions directly from original client materials, without prior expert reformulation? Existing OR benchmarks commonly begin with self-contained problem descriptions or predefined optimization tasks, leaving the work of translating original business materials into professionally validated solutions underrepresented. We introduce OR-License, an end-to-end benchmark of 31 instances across 10 industrial scenarios drawn from paid OR engagements totaling $3.68 million in project value. It preserves original requirements, input data, deliverable formats, and information boundaries after de-identification. To evaluate information acquisition throughout optimization, we further introduce ORSim, a controlled client simulator that discloses business facts without supplying OR expertise. Hidden evaluation checks feasibility and compares objective quality against original expert solutions accepted by clients before benchmark construction. Deployment-Adjusted Quality (DAQ) averages expert-relative quality across instances, assigning zero credit to infeasible deliveries. Across 11 frontier LLMs evaluated with coding-agent harnesses, several exceed the expert reference in conditional objective quality, but none matches its DAQ; the strongest system achieves 83.87% strict feasibility and a DAQ of 0.84 against the expert reference of 1. Client information acquisition extends beyond initial clarification into solution development and validation, with some later stages exhibiting higher confirmed issue-resolution proportions. Trajectory analysis identifies failures spanning formulation and implementation, while task-equalized research-chain completion rates range from 24% to 45% across models on evaluable tasks. These findings motivate sustained information-gap resolution, faithful problem formulation, and evidence-linked self-improvement for reliable long-horizon agents.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.