ShopLife-Bench: A Benchmark for Multi-Session Multi-User Agent Interaction in E-Commerce Environments
Abstract
Existing benchmarks for conversational AI agents mostly simulate a single user in a single session. They do not capture multi-user state management, the dynamic evolution of user behavior, or cross-session task decomposition and feedback. As a result, they cannot cover the complex, path-dependent behavior of real users across sessions. We propose ShopLife-Bench, a benchmark for multi-session, multi-user agent interaction in e-commerce environments, with four key contributions: 1) Our user simulator does not model internal mental states. Instead, it distills users' reactions to agent errors and their exit patterns from over ten thousand real e-commerce conversations, and builds a reproducible behavior policy driven by accumulated errors. 2) We restructure user profiles into structured, addressable atomic entries, so that every addition, update, and deletion can be verified item by item. We also design tasks in which the current request alone cannot uniquely determine the target product; the correct choice requires profiles maintained from earlier sessions. These tasks jointly evaluate profile maintenance and personalized use. 3) Cross-session tasks are built around the same user's product decisions and fall into three types: dependent, independent, and interleaved. Explicit statements, implicit needs, and noise are added on top. 4) Each evaluation session includes two task users and one distractor user who only asks about products or chats. This setting tests the agent's ability to attribute identities, switch between tasks, and isolate user states. Experiments on six representative models show that after shifting to multi-user concurrency, profile recall and order accuracy generally drop; even the best model reaches only order accuracy and a zero-error rate. Cross-user state isolation and cross-session consistency remain critical bottlenecks for current agents. Code and evaluation data are available at https://anonymous.4open.science/r/ShopLife-Bench.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.