acceptodds
Under review as a conference paper at ICLR 2027

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios

Abstract

Large Language Models (LLMs) are increasingly deployed as agents that interact with stateful environments over multiple steps: gathering hidden information, composing tool calls, and committing state changes. We refer to this capability as **multi-step tool use**. Existing benchmarks have advanced tool-use agent evaluation, but often focus on isolated API calls, short trajectories, or settings that are difficult to scale or control. We introduce , a ***fully synthetic*** benchmark with state-changing tasks across three product domains: *Honor of Kings*, *QQ Music*, and *Tencent Meeting*. decouples *environment synthesis* from *task synthesis*: graph-guided database filling builds reusable, orphan-free product environments, while generator-solver asymmetry creates tasks with both an *information gap* and a *tool gap*, requiring agents to discover hidden data and compose multiple tool calls before changing state. Outcomes are graded deterministically by database-state diffs. Since both environments and tasks are synthetic, is controllable at the environment level and scalable at the task level. Benchmarking cutting-edge LLMs shows that multi-step tool use remains challenging: Pass stays below for the strongest models, and even with code execution in -Code extension, reliability (Pass) remains below .

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.