acceptodds
Under review as a conference paper at ICLR 2027

Beyond Semantic Matching: Evidence-Grounded Multi-Agent Product Deep Search

Abstract

Modern product search excels at semantic matching, but realistic shopping queries often demand complex, multi-condition requirements—spanning appear- ance, specifications, compatibility, price, and user reviews—that holistic relevance scores alone cannot reliably guarantee. We introduce Product Deep Search, a benchmark and paradigm for evidence-grounded product retrieval. Candidate products are represented through structured product-page evidence, and models use tools to retrieve supporting or contradictory evidence across heterogeneous sources. To solve this problem, we propose PDeepSearch, a multi-agent frame- work that turns semantic search into evidence-grounded decision making. Start- ing from candidates produced by multimodal retrieval and reranking, PInspect independently examines each product with evidence-access tools and constructs condition-level profiles of support, uncertainty, and contradiction; PDecide then compares these profiles across candidates to make the final selection. The frame- work can be instantiated with both API-based and open-weight models, and con- sistently improves strong zero-shot backbones by 2.43–7.29 Hit@1 points. We further propose PDeepSearch-9B, a product search agent trained by synthesiz- ing inspection and decision trajectories. PDeepSearch-9B achieves 69.43 Hit@1, outperforming larger open-weight models, including Qwen3.5-27B (63.71) and Gemma-3-27B-IT (62.86), while approaching competitive top-tier models like Gemini-3.5-Flash and Kimi-K2.6. Further analysis shows that evidence verifica- tion acts as an effective correction layer over semantic reranking, with the largest absolute gains observed on queries combining a reference image with visual, de- tail, and review-grounded conditions. These results highlight evidence-grounded verification as a key capability beyond semantic matching for complex product search, and show that it can be effectively learned by compact models.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.