SearchPLM: Towards Search-Proficient Language Models
Abstract
Despite recent advances in reasoning and tool use, Large Language Models (LLMs) still struggle with complex information-seeking tasks. Through systematic analysis, we identify a key failure mode in current LLM-based search agents: they tend to issue overly specific queries that closely overlap with the original question and terminate search prematurely after several retrieval failures. We term this phenomenon the Golden Document Fallacy, an implicit assumption that a single document can fully answer the original question , which reduces multi-turn search to a shortcut-like search process. To address this issue, we first collect, identify, and summarize search experience from human experts, and distill it into a general Exploration-Coverage-Synthesis (E-C-S) paradigm, which structures effective search into strategic exploration, iterative evidence coverage, and multi-source synthesis. Based on this paradigm, we develop SearchPLM, a Search-Proficient Language Model trained with strategy-guided supervised distillation and curriculum-based reinforcement alignment. Experiments show that SearchPLM outperforms strong open-source baselines with more parameters, highlighting the importance of search strategy optimization for more generalizable information-seeking tasks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.