OR-Space: A Workspace-Based Benchmark for Industrial Optimization Agents
Abstract
Large language model (LLM) agents are increasingly used for operations research (OR) modeling, but existing OR benchmarks mostly test model construction from a self-contained problem statement. Industrial optimization also requires connecting requirements to data and code, revising models, and explaining solver results. We introduce OR-Space, a benchmark of executable workspaces containing business documents, structured data, code, solver outputs, and task-specific evaluators. OR-Space evaluates three tasks separately: BUILD, where agents create solver-ready models from heterogeneous artifacts; REVISE, where they update models under changed requirements while preserving valid prior logic; and EXPLAIN, where they ground answers about solutions and constraints in workspace evidence. OR-Space evaluates 16 language models on 100 base problems instantiated as 300 task instances, revealing substantial failures in schema mapping, constraint grounding, legacy-code use, and solver-backed explanation. Controlled studies show that workspace interface, revision context, and solver backend affect performance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.