acceptodds
Under review as a conference paper at ICLR 2027

Counterfactual Edit Credit with Semantic Retrieval Feedback for Product Query Rewriting

Abstract

Product queries are often short and underspecified, leaving much of the user's preference implicit. Recent work trains large language models (LLMs) as query rewriters with reinforcement learning (RL) using ranking rewards based on the ranks of labeled relevant products. However, these rewards are sparse when relevant products are not retrieved and coarse when different retrieval outcomes leave their ranks unchanged. We propose Counterfactual Generative Search (CGS), an RL framework that uses dense semantic retrieval feedback to assign counterfactual credit to individual query edits, while retaining ranking reward as the objective for the complete rewrite. uses a relevance teacher to score retrieved products against available evidence about user intent and a semantic list utility to distinguish retrieval outcomes beyond labeled-product ranks. It assigns credit to each edit by comparing retrieval quality when the edit is removed from the complete rewrite and when it is applied to the source query on its own. With extensive experiments on H&M and Amazon ESCI, we show that CGS consistently outperforms baselines, improving NDCG@100 by 10.3% and 4.9%, respectively.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.