acceptodds
Under review as a conference paper at ICLR 2027

Distill What You Retrieve: Candidate-Set On-Policy Distillation with Reliable Preference Correction

Abstract

Dense retrieval is essential for LLM agents to identify relevant tools and reusable skills from large candidate collections. Knowledge distillation (KD) has become a widely adopted paradigm for transferring the capabilities of powerful teachers to compact student retrievers. However, conventional offline distillation relies on static candidate pools that cannot adapt to the student’s evolving retrieval behavior. On-policy distillation (OPD) addresses this mismatch by collecting teacher supervision on student-induced candidates, yet its application to dense retrieval remains limited by two intrinsic issues: query-candidate-level supervision fails to capture the setwise competition that governs retrieval, and soft distribution matching does not explicitly guarantee that annotated positives outrank negatives. To address these issues, we propose a two-stage framework that combines candidate-set on-policy distillation with targeted preference optimization. In Stage 1, Candidate-Set On-Policy Distillation (CS-OPD) constructs query-specific Top-K candidate sets from the training student and transfers the teacher’s listwise preferences over the same candidates via KL divergence. In Stage 2, Reliable-Disagreement Direct Preference Optimization (RD-DPO) identifies reliable residual ranking errors through teacher-validated margin disagreement and corrects them using a score-level DPO objective tailored to dense retrieval. Extensive experiments on ToolRet and SkillRet demonstrate consistent improvements over strong retrieval and distillation baselines. Further analyses confirm that CS-OPD improves candidate-set distribution alignment, whereas RD-DPO resolves residual ordering inconsistencies, highlighting their complementary contributions.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.