Searching in the Dark: How Do Search-Augmented LLMs Handle Abstract Queries?
Abstract
Current search benchmarks largely focus on well-specified, multi-constraint queries that determine a target answer. However, in practice, users often omit constraints and may describe what they are looking for indirectly. We study abstract search: queries that leave multiple plausible answers, and whose available constraints may be expressed explicitly or described indirectly. We construct a dataset of 515 queries spanning two search settings: 255 human-authored Entity Recall queries, where a user has a target entity in mind but provides partial or indirect cues that could fit other candidates and 260 Open Discovery queries, where users state their criteria directly but those criteria include multiple valid candidates. Across seven search-augmented LLMs, we find that models struggle to balance retrieval breadth and candidate validity: most models either retrieve too conservatively and miss valid candidates or retrieve too broadly and admit invalid candidates, with no model achieving both high precision and coverage. We observe that models struggle to resolve indirect descriptions in underspecified queries, with the strongest model recovering only 67.5% of user-intended entities in Entity Recall and indirect descriptions are associated with the largest drop in retrieval performance. Models also remain overconfident: Web search improves calibration relative to parametric memory, but this improvement comes from accuracy rising towards confidence, not from models becoming more cautious. Overall, our findings highlight the need to characterise search behaviour beyond well-specified settings and reveal important blind spots that emerge when users search abstractly.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.