CLEAR: Learning to Clarify for Earlier Gains in Domain-Specific Interactive Retrieval
Abstract
Domain-specific document retrieval often begins with queries that leave relevance-critical information unspecified. Interactive retrieval elicits this information through clarification questions and query updates. However, clarification must adapt to the evolving dialogue, and questions must be prioritized to improve retrieval with fewer rounds. We propose Clarification Learning for Early and Accurate Retrieval (CLEAR), a framework that uses retrieval results and dialogue history to identify what to clarify, and retrieval feedback to learn question priorities. Progressive Residual Facet Discovery (PRFD) uses residual quantization to organize document contrasts, combining them with domain facts and dialogue history to identify requirement facets, or aspects of the information need that can be clarified. PRFD produces candidate facets; the policy selects a facet and generates a question. Turn-aware Hierarchical Policy Optimization (THPO) constructs a return matrix from per-round retrieval gains and uses hierarchical comparisons across facet categories and candidate facets for reinforcement learning. Across legal, medical, and patent benchmarks, the domain-specific configurations achieve higher NDCG@10 at R4 and mean NDCG@10 over R1-R4 than the compared interactive baselines. The patent configuration disables round-dependent prefix selection, assessing framework transfer rather than the full progressive mechanism. Legal and medical trajectories additionally quantify the proportion of queries that first reach an NDCG@10 threshold within each round budget.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.