acceptodds
Under review as a conference paper at ICLR 2027

CommerceAgentBench: Benchmarking Long-Horizon Agents in High-Fidelity, Stateful, and Reproducible Replicas of Real Online Services

Abstract

Whether large language model agents can carry out sustained operational work remains a central evaluation challenge. Such work spans tools and systems, changes shared state, and depends on repeated judgment. CommerceAgentBench (CAB) evaluates this ability in high-fidelity, stateful, reproducible replicas of real services across browser UI, API/MCP, isolated command-line, replay, and open-source appliance forms. Task-surface fidelity aligns the declared routes, controls, validations, visible outputs, and state transitions with upstream evidence and regression tests. We evaluate a fixed set of 107 cases spanning commerce, sourcing, office suites, marketplace operations, logistics, and ERP-like reconciliation. Human authors reconstruct workflows from request examples or curated procedures, ground workspaces in public or officially sourced synthetic records, and define outcome contracts. Agent actions mutate durable state, and verifiers read final state or artifacts. Scoring combines strict workflow completion with macro-averaged capacity over 834 checks: 799 non-judge checks and 35 grounded boolean model-judge checks. Thirteen models run on three agent frameworks with three independent runs per model–framework pair. The strongest model achieves a 59.7% mean pass rate and 0.892 capacity across frameworks. On a selected 15-cell diagnostic panel, supplying object-selection decisions or target-system state enables previously failed workflows to complete. A 14-verifier state audit and one-change replays on 12 of 32 HARD cases test score responses to controlled errors. These results characterize complete workflows, partial progress, and the effect of targeted interventions. Our data and code are available at https://anonymous.4open.science/r/CommerceAgentBench-5E42.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.