acceptodds
Under review as a conference paper at ICLR 2027

ORWorkerBench: Benchmarking LLMs from Data to Decisions under Uncertainty

Abstract

Operations Research increasingly provides the decision layer for AI systems that must allocate resources, schedule work, manage inventory, or act under uncertainty. Recent optimization benchmarks show strong performance from frontier LLMs and agents, but many begin from a specified optimization problem, where the stochastic quantities required by the model are already provided. In real operational settings, those quantities often have to be inferred from finite observations before optimization can begin. We introduce ORWorkerBench, a benchmark for evaluating this observation-to-decision setting. Each underlying decision problem is evaluated under matched conditions: in NL2Code, the relevant stochastic parameters are given explicitly, while in Workflow, the agent must infer them from 100 observations before constructing and executing the optimization procedure. Our benchmark contains 147 NL2Code tasks and 441 Workflow tasks across 21 human-verified executable protocols. We additionally use neuron-guided synthesis to reduce behavioral redundancy during benchmark construction; compared with embedding-based selection, it produces more diverse activation patterns across four held-out models. Our evaluation of six proprietary and two open-weight LLMs reveals a substantial observation-to-decision gap. The strongest models achieve over 99% tolerance accuracy when the stochastic quantities are supplied, yet incur roughly 7% mean capped error when those quantities must instead be recovered from observations. Moreover, even the strongest systems perform close to a simple same-data reference that estimates the required stochastic quantities from the same observations and passes them to a verified solver. These results show that strong performance on specified optimization problems does not yet translate into reliable decision making from finite observations.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.