acceptodds
Under review as a conference paper at ICLR 2027

ImmuneCoT: Learning Reasoning Immunity via Privileged On-Policy Self-Distillation

Abstract

Chain-of-thought (CoT) reasoning helps large language models (LLMs) solve complex problems, yet unsafe intermediate steps can steer subsequent reasoning toward harmful assistance. Learning to recover from such drift is therefore an important component of reasoning safety. We introduce ImmuneCoT, an immune-inspired on-policy self-distillation (OPSD) framework that trains recovery through controlled exposure to corrupted reasoning prefixes. A frozen teacher initialized from the same base model provides two complementary supervision signals: Recognition conditions on the infected prefix to elicit drift-aware correction, whereas Response bypasses that prefix and anchors behavior to the original request. A Base-adjusted product of experts fuses these signals into token-level targets for the student's on-policy continuations, enabling recovery without privileged triggers at inference. Auxiliary benign examples provide self-tolerance training to limit indiscriminate refusal. We evaluate Qwen3-1.7B and Qwen3-4B on standard safety, over-refusal, and general-capability benchmarks, and assess Qwen3-4B under H-CoT mid-reasoning hijacking. Among cases with safe answers before injection, ImmuneCoT achieves a safe recovery rate of 55.5%; on the common Clean-safe subset shared with STAR-1, its paired advantage is 26.1 percentage points. On Qwen3-4B, it also reduces HarmBench attack success from 30.9% to 6.2%, keeps over-refusal below 4% on both benign benchmarks, and limits the maximum observed capability decrease to 0.4 percentage points relative to Base. These findings support learning recovery from corrupted reasoning states while retaining helpfulness and general capability. Code, data, and model resources are available at https://anonymous.4open.science/r/ImmuneCoT-3DE0/.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.