ReSCC: Calibrating Agent Self-Distillation with Retrieved Action Outcomes
Abstract
Teacher-guided self-distillation provides long-horizon agents with dense token- level supervision, complementing the sparse and coarse feedback from terminal rewards. Confidence-based gating reflects how strongly a teacher prefers a pre- diction but does not establish whether the resulting action improves downstream outcomes. We propose RESCC (Retrieval-Simulated Counterfactual Credit), which calibrates teacher guidance using action outcomes from previously collected training trajectories. RESCC retrieves similar state–action histories and compares the current action with well-supported alternatives using behavior-probability correction and a cross-fitted outcome model. It converts this evidence into a bounded signed weight on the distillation loss and falls back to the original update when historical support is insufficient. GRPO and online rollouts remain unchanged. Across ALFWorld, WebShop, and Search-QA with Qwen2.5-3B, Qwen2.5-7B, and Qwen3-1.7B, RESCC achieves improvements of up to 19.5 percentage points over existing baselines. Ablation studies show that outcome-aware retrieval, behavior correction, support-aware abstention, and bidirectional credit all contribute to the performance gains. Executed-alternative evaluations further show that RESCC aligns substantially better with realized action outcomes than confidence-based signals, with reliability improving as historical support increases.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.