Memory Is Not Control: Recovering Explicit Forgetting in Reasoning-Distilled Models via Contrastive Routing
Abstract
Large language models (LLMs) have demonstrated strong reasoning abilities, which distillation transfers to smaller models. However, reasoning-distilled models recall earlier information better than their base models on a multi-turn dialogue benchmark while following earlier instructions less reliably, illustrating that memory does not ensure control. In this paper, we study this gap through explicit forgetting and find that reasoning-distilled models achieve higher immediate forgetting compliance than their paired base models, but this advantage shrinks or reverses on delayed probes. In contrast, base models often avoid forgotten content that the distilled models reuse. Yet injecting base-model activations or logits during generation fixes some answers while breaking others, since it acts before the original answer exists. Inspired by these findings, we propose Forgotten-State Contrastive Routing (FSCR), which turns the forgotten-content leakage of both completed answers into a contrastive routing logit. Concretely, FSCR switches to the base answer only when this logit is non-negative. The logit is the minimum of three margins computed from the lexical leakage of each answer and the gap between them. We evaluate FSCR on multiple explicit-forgetting benchmarks with several reasoning-distilled models, illustrating consistent improvements with both annotated and automatically extracted forgotten content, demonstrating that FSCR can achieve better forgetting compliance across models and dialogues. The gains are largest after intervening, showing the effectiveness and robustness of our analysis findings and the proposed FSCR. Our code and data are available at https://anonymous.4open.science/r/fscr-review-DEF8
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.