RecoAtlas: A Benchmark and RL Environment for LLM Recommendation Agents Optimizing Set-Level Utility
Abstract
LLM recommender agents increasingly produce structured reports: sets of items accompanied by natural-language justifications. Yet existing benchmarks often focus on reranking small candidate sets or judging reports primarily by semantic plausibility, providing limited support for evaluating and training agents based on set-level utility. We introduce Recommender Atlas (Agentic Tasks for Learning and Assessment in Shopping), or RecoAtlas, a benchmark and reinforcement learning environment grounded in user behavior. RecoAtlas provides a shopping toolkit for product retrieval, complementarity, and diversity. Given a shopping query, an agent uses these tools to generate a structured recommendation report and receives behavior-grounded feedback, forming an end-to-end interaction that supports both evaluation and reinforcement learning. Its controlled tool environment exposes agents to semantic, behavior-aligned, or faulty tools, enabling systematic analysis of how LLM reasoning, tool signals, and tool-use policies contribute to performance. Across controlled experiments, RecoAtlas exhibits key properties of a meaningful benchmark for agentic systems: performance scales with model capacity and test-time compute, and improves with stronger tools. These experiments also reveal that semantic plausibility does not necessarily translate into behavior-grounded utility. RecoAtlas provides a foundation for training and evaluating shopping assistants to move beyond semantic plausibility toward coherent, behaviorally grounded recommendation sets.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.