See Less, Shop Better: Efficient Visual Evidence Acquisition for E-commerce Search Agents
Abstract
E-commerce search agents are increasingly used to find and recommend products based on complex user requirements. However, many decision-critical product attributes are available only in image-heavy detail pages. Directly exposing all visual content from a large set of products to an agent is both computationally expensive and difficult to scale. To address this challenge, we study how e-commerce search agents can effectively and efficiently utilize visual product information. We first introduce two complementary evaluation tasks: Fixed-Item Question Answering, which isolates and directly evaluates an agent's ability to comprehend multimodal product details, and Open-Ended Product Recommendation, which holistically assesses its end-to-end performance across the full search-and-recommendation pipeline. We then decouple product visual understanding from the main agent through a visual evidence tool, allowing the agent to selectively acquire visual information on demand. Building on this framework, we develop a post-training pipeline combining a staged Supervised Fine-Tuning (SFT) curriculum with Reinforcement Learning (RL) under constrained tool-turn budgets. Experiments show that our staged SFT substantially boosts both multimodal product-detail understanding and open-ended recommendation, while budget-constrained RL further improves the efficiency–effectiveness trade-off—enabling our 9B agent to achieve a 46.8% recommendation quality rate under a strict 4-turn budget, surpassing GPT-5.6 Sol's 45.6% with unconstrained tool use.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.