AlgoWorlds: Benchmarking Tool Use for Global Optimization in Algorithmic Worlds
Abstract
Tool-use benchmarks generally evaluate whether an agent completes a workflow using appropriate tools and valid arguments. However, feasibility alone is insufficient in real-world decision settings such as route planning and fleet dispatch. Individual choices interact through shared constraints and costs, so a feasible solution may still be substantially suboptimal. This raises a harder question: can an agent turn information gathered through tools into a globally optimal decision? In this paper, we introduce AlgoWorlds, a benchmark that transforms formally specified combinatorial optimization problems into partially observed decision environments with global optima. Each environment contains a hidden optimization instance that the agent observes only through sequential calls to task-specific information tools. The agent then commits to one structured decision, which an independent checker evaluates for feasibility, objective value, and optimality. AlgoWorlds contains 240 such environments, covering ten combinatorial optimization families and four workload levels. To make these environments both controllable and auditable, family-specific deterministic programs generate hidden instances, while exact algorithms certify their global optima and determine their workload levels. Each instance is exposed through two structurally different multi-tool interfaces, which preserve the same decision problem under different information presentation. We evaluate seven leading LLMs, including Claude Opus 4.8 and GPT-5.6 Sol, on AlgoWorlds. Results show that achieving global optimality remains highly challenging: although leading models achieve feasibility in most cases, the best-performing model achieves exact optimality in only 38.61% of cases. Further analysis shows that even when the LLM agents collect sufficient information to reconstruct the hidden instance, most failures end in feasible but suboptimal decisions. The challenge therefore extends beyond information acquisition to information integration, global constraint reasoning, and decision verification.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.