acceptodds
Under review as a conference paper at ICLR 2027

Information-Gain Calibrated On-Policy Distillation for Deep Search Agents

Abstract

On-policy distillation (OPD) has recently emerged as a promising approach for transferring the capabilities of large agents to compact students. In deep search, where agents iteratively retrieve external evidence, a teacher's preference among plausible queries before retrieval may not predict which queries will return useful evidence. Standard OPD nevertheless supervises each student query from the preceding context alone, inducing a failure mode we term evidence-blind supervision (EBS). To understand this, we regroup the OPD objective by interaction turns and reveal two implicit priors: a budget prior that allocates supervision by token volume, and a direction prior that favors tokens preferred by the teacher, neither of which accounts for the evidence returned by the current query. Our empirical analysis shows that both priors contribute to EBS in practice: the first search turn accounts for 51.1% of total absolute information gain but receives only 18.0% of the supervision budget, while 43.7% of search queries have teacher rewards that disagree in sign with post-retrieval information utility. To mitigate this, we propose \ours, which uses post-retrieval information gain (IG) to calibrate both priors: reallocating supervision toward information-critical turns while calibrating query preferences using the utility of retrieved evidence. Experiments on six multimodal search benchmarks show that \ours improves average accuracy by 3.4 and 1.7 percentage points for 2B and 4B students over standard OPD, enabling the 4B student to surpass its 8B teacher on MMSearch and BrowseComp-VL.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.