acceptodds
Under review as a conference paper at ICLR 2027

OL-Bench: A Comprehensive Benchmark for Evaluating AI Agents in Omni-Life Scenarios

Abstract

Large language model (LLM)-based agents are increasingly applied to real-life interactive scenarios, where they are required to understand evolving user needs, follow environmental constraints, interact over multiple turns, and use tools to complete tasks. However, existing benchmarks often provide limited scenario coverage, simplify interactive environments, and primarily evaluate final task completion, creating a significant gap between benchmark evaluation and real-life application settings. Consequently, they may fail to faithfully reflect agents' actual performance in practice. In this work, we present Omni-Life Benchmark (OL-Bench), a comprehensive interactive benchmark for evaluating agents across diverse real-life scenarios. OL-Bench adopts an Entity–Tool–Instruction hierarchy to synthesize executable interactive environments. Based on these environments, dependency-aware multi-turn tasks are constructed with explicitly planned answers and multi-level distractors across four difficulty levels, ranging from easy to extreme. These tasks are executed through a scenario-specific Harness that coordinates environment state transitions, tool use, and interaction trajectories. OL-Bench contains 1055 tasks across 218 real-life scenarios, with an average of 46.4 tools and 20.2 Skills per scenario. It further provides 18 metrics across seven capability dimensions to evaluate both task completion and interaction processes. Our comprehensive evaluation shows that GPT-5.6-Sol, the best-performing model, achieves a Full Task Completion Rate of only 18.0%, underscoring the difficulty of complex real-life interactions for current agents. These results demonstrate the value of OL-Bench for systematically evaluating agent capabilities and diagnosing their limitations in real-life interactive scenarios.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.