acceptodds
Under review as a conference paper at ICLR 2027

Learning When Evidence-Sensitive Guidance Helps: Calibrating Counterfactual On-Policy Distillation for Memory-Grounded Question Answering

Abstract

Memory-grounded question answering requires language models to use supporting evidence from past interactions. Counterfactual on-policy distillation provides evidence-sensitive supervision by contrasting teacher predictions with and without that evidence. However, the contrast alone does not determine how strongly the teacher target should be adjusted. We propose Mirror-Calibrated Counterfactual On-Policy Distillation (MCC-OPD), which uses a mirror KL budget to adapt the strength of evidence-sensitive teacher corrections. At each student-generated prefix, the method retains the full-context teacher distribution as its base and limits the counterfactual correction by the KL divergence of an equal-strength move in the opposite direction. It selects the largest feasible correction coefficient and trains the student with reverse KL, without answer-level rewards or a learned value model. We establish convexity of the target divergence, enabling a scalar search for the calibrated coefficient. We also analyze outcome-based calibration through local student utility and compare the two calibration strategies empirically. Trained on only 140 LoCoMo questions, MCC-OPD outperforms strong baselines including full-context OPD, CROP, SA-OPD, and ExOPD in LLM-judge accuracy across LoCoMo, REALTALK, MuSiQue, and LongMemEval, with the largest absolute gain of 6.54 percentage points over CROP on LoCoMo. Matched ablations further show that calibration improves LLM-judge accuracy over uncalibrated counterfactual distillation on all four benchmarks.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.