acceptodds
Under review as a conference paper at ICLR 2027

BizBench: Unveiling Hidden Challenges in Deploying Long-Horizon Agents for Real-World Business

Abstract

Existing evaluations of long-horizon agents often emphasize final task success while overlooking the validity of intermediate decisions. Real-world business applications, however, demand stricter assessment: an agent must behave correctly not only at the endpoint, but throughout the entire execution process. This requirement is particularly critical in high-stakes domains such as payments, where invalid intermediate actions can cause financial loss, regulatory violations, or irreversible downstream consequences even when the final goal is eventually achieved. Consequently, agents that perform well on existing benchmarks may still fail in deployment, as outcome-centric evaluation provides limited insight into the validity and downstream consequences of intermediate decisions. In this work, we aim to narrow the gap between offline evaluation of long-horizon agents and their deployment in real-world business systems. We introduce BizBench, a benchmark derived from real-world business procedures at a leading payments company. We encode these procedures in an executable environment with aligned policies, states, and tools; compose long-horizon tasks with resource, temporal, and commitment constraints, requiring agents to reconcile knowledge conflicts and identify outdated guidance; and replay reference traces to verify joint feasibility of goals and constraints. The resulting benchmark comprises 100 executable tasks and evaluates both final task success and execution validity through deterministic checks over outcomes and execution histories, while permitting multiple legitimate solution strategies. Extensive experiments on BizBench show that even state-of-the-art models struggle on BizBench, with some runs extending beyond 200 steps of tool use and document retrieval. Our results reveal a substantial gap between successful task completion and sustained valid execution. BizBench provides a testbed for developing long-horizon agents that can coordinate information gathering, planning, and action to complete business tasks while maintaining valid execution throughout.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.