acceptodds
Under review as a conference paper at ICLR 2027

FashionArena: A Large-Scale, Production-Grounded Environment for Training and Evaluating Fashion Shopping Agents

Abstract

Fashion shopping agents are increasingly built to identify user intent and find products with multimodal inputs. Existing fashion shopping environments mainly face two limitations: (i) insufficient product scale and (ii) reliance on catalog-based retrieval, which can produce discrepancies between offline and online tool feedback. These limitations may interfere with agents’ learning and exploration. We introduce FashionArena, a large-scale, production-grounded environment and benchmark for training and evaluating fashion shopping agents. It provides approximately 1.6 million real products and six executable shopping tools, with online-concordant feedback grounded in millions of cached query–response records. Our tasks are constructed from real user intents captured in online behavior logs, covering text-only and image–text requests. We categorize these intents by their shopping needs and constraints, then collect candidate products and agent trajectories through online tools. Human annotation assesses both product suitability and the quality of tool use throughout agent trajectories. The resulting dataset contains 20K training tasks and 2K held-out evaluation tasks. Evaluation follows human-defined criteria and accepts suitable products beyond the reference answers. Experiments show that offline task performance closely matches online task performance. Supervised fine-tuning followed by reinforcement learning significantly improves recommendation performance on our benchmark. These gains persist when the trained policies interact directly with online tools. We release the environment, datasets, and training code to support reproducible research on fashion shopping agents.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.