How Models Reason Matters: Task Circuit Replay for Continual Instruction Tuning
Abstract
Mitigating remains a major challenge in continual instruction tuning of large language models. Existing methods preserve past knowledge mainly through historical examples, model outputs, or selected internal representations. However, we argue that effective knowledge preservation should go beyond retaining such past information and instead preserve the underlying capabilities acquired from previous tasks. This shifts the focus toward how these capabilities are represented in the model. We argue that how the model computes for a task provides an important characterization of learned capabilities, as task-level computation patterns can consistently support task behavior across different inputs.Based on this perspective, we propose , a continual instruction-tuning framework that represents and preserves learned capabilities through task-relevant circuits. TCR identifies sparse task-level circuits spanning different model layers, and retains their characteristic responses together with a small number of historical examples. When learning subsequent tasks, these examples reactivate the corresponding circuits, whose responses are selectively preserved to maintain the capabilities they support. To facilitate task-specific circuit identification, TCR introduces Task-BOS to explicitly condition model computation. In this way, TCR preserves past knowledge as reusable task-level computation patterns rather than merely revisiting past observations. Experiments on TRACE and SuperNI demonstrate that TCR consistently reduces forgetting while improving average accuracy and backward transfer.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.