SOCAL: Step-Specific Over-Search Calibration for Search Agents via Self-Distilled Hindsight
Abstract
Search agents are prone to over-search when they fail to accurately assess whether their internal knowledge is sufficient, resulting in unnecessary external search. Outcome-based RL can further amplify this miscalibration, as trajectory-level credit does not distinguish between useful and unnecessary search decisions. Existing process-aware methods incorporate process-level feedback on search necessity, yet the degree of correction applied to individual search decisions remains largely determined by predefined reward formulations or advantage-rescaling rules. This raises a natural question: can the agent itself provide an adaptive, decision-specific signal for calibrating its search behavior? To this end, we introduce SOCAL (Step-Specific Over-Search Calibration for Search Agents via Self-Distilled Hindsight), which transforms over-search diagnoses into step-specific hindsight and leverages self-distillation for adaptive search calibration. Specifically, SOCAL first performs knowledge probing over sampled trajectories to extract hindsight reflecting the model's knowledge sufficiency. For each diagnosed step, it contrasts the policy's support for the same sampled step under the original and hindsight-conditioned views. The resulting likelihood discrepancy serves as a self-distilled signal for reshaping the corresponding outcome-derived credit, allowing the correction strength to adapt to each individual search step. Extensive experiments on seven knowledge-intensive question-answering benchmarks demonstrate the effectiveness of SOCAL, yielding 1.5–2.9 percentage-point gains in CEM over the strongest baselines and the lowest over-search rates across both 3B and 7B model settings. Our code is publicly available at https://anonymous.4open.science/r/SOCAL.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.