acceptodds
Under review as a conference paper at ICLR 2027

ORC-Bench: Benchmarking LLMs Capabilities for Optimization and Reasoning Under Constraints

Abstract

Large Language Models (LLMs) have achieved great performance on a broad range of knowledge and reasoning benchmarks, yet existing evaluations rarely probe their ability to reason and optimize end-to-end under interacting structural, physical, and operational constraints. We introduce ORC-Bench, a benchmark for evaluating LLMs on Optimization and Reasoning under Constraints over realistic structured inputs. ORC-Bench comprises 30 tasks spanning ten categories, three difficulty levels, and two data modalities, jointly probing six skills: graph understanding, tabular understanding, mathematical formulation, arithmetic execution, problem abstraction, and multi-step optimization. Unlike existing benchmarks, ORC-Bench scores models with constraint-aware metrics derived from physics-informed and domain-specific checkers, going beyond accuracy to measure structured-output validity and constraint-satisfaction rates. Evaluating nine LLMs across architectures, scales, and reasoning regimes, we find that while scale and explicit reasoning improve performance on our tasks, all models (including frontier reasoning models with human performance on MATH-500 and ARC-AGI) struggle to satisfy interacting physical and logical constraints while optimizing domain specific objectives.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.