Confidence-Guided Query Refinement for Search-Augmented Reasoning
Abstract
Search-augmented reasoning enables language-model agents to acquire missing knowledge, but retrieval can fail when an agent generates a query that poorly expresses its current information need. We study this failure mode, which we term query unreliability, and show that token-level query confidence provides a pre-retrieval signal of retrieval quality and downstream answer correctness. Based on this observation, we introduce Confidence-Guided Query Refinement (QCQR), a post-hoc test-time method that executes confident queries unchanged while selectively rewriting, decomposing, or disambiguating low-confidence queries using an off-the-shelf reformulator. QCQR requires no additional training of the search agent or retriever. Across heterogeneous search agents and knowledge-intensive QA benchmarks, QCQR improves average answer accuracy by up to up to 6.18 percentage points, improves gold-document retrieval and downstream calibration, and requires a comparable number of external search calls. Unconditional and random refinement do not yield consistent gains, highlighting the importance of selectively intervening on uncertain queries. Our results establish query confidence as a practical signal for repairing unreliable search actions before retrieval.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.