Seek What You Ignorant: Boundary-Aware Behavioural Calibration for Search Agents
Abstract
Reliable agentic search requires large language models to calibrate trust in paramet- ric memory and retrieved evidence, deciding when to use either source, integrate both, or abstain. Yet outcome rewards, whether binary or graded, provide only sparse feedback on completed trajectories and limited direct guidance for these intermediate decisions. Moreover, appropriate reasoning patterns vary across knowledge states. We propose Knowledge boundary Self-Distillation (KbSD), a quadrant-adaptive self-distillation framework that combines dense token-level supervision with GRPO-based reinforcement learning. KbSD estimates parametric certainty and response stability from repeated retrieval-free samples, and retrieval quality from the relevance of retrieved passages. These signals, together with quadrant labels and ground-truth answers, form privileged hints for a teacher that shares the student’s model parameters. Conditioned on these hints and the student’s observed history, the teacher provides token-level guidance to the hint- free student. This information asymmetry enables dense supervision without a larger or separately trained teacher, and the student requires no privileged hints at inference. To accommodate differences in reasoning diversity across the four knowledge-boundary quadrants, KbSD uses mode-seeking reverse KL for concen- trated evidence integration, mass-covering forward KL for diverse refusal patterns, and Pareto-optimal bidirectional KL for states that require both precise source selection and broad coverage. Experiments across nine benchmarks show lower unreliability than the evaluated baselines. We report task F1 alongside refusal rates to characterize the accuracy–refusal trade-off and interpret reliability gains in the context of answer coverage.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.