CANARY: Separating Behavioral Anchoring from Risk-Guided Selection in Continual LLM Post-Training
Abstract
Continual post-training often couples a retention objective with a policy for selecting the behaviors it protects, making their contributions difficult to separate. CANARY distinguishes cached stage-start response anchoring from risk-guided anchor allocation. Across six matched seed blocks with Qwen2.5-3B on a fixed tool-use, chemistry, and medical stream, gradient-ranked anchoring improves final stream performance, forgetting, and held-out accuracy over the implemented live-teacher on-policy self-distillation baseline in every block, with mean paired gains of , , and , respectively. Yet scored random and bottom-ranked anchors achieve stronger aggregate results. Across six saved states from one seed, gradient-conflict scores weakly agree with cloned one-microbatch AdamW loss changes, with mean Spearman correlation . Ranking by the cloned update yields favorable mean effects relative to both matched selector controls, with joint non-worsening in three of six runs against each. A four-seed risk-screened coverage follow-up improves mean held-out accuracy over random selection but worsens stream performance and forgetting. A separate supervised-replay comparison finds no benefit from adding the tested anchoring branch. These results distinguish the benefit of an anchored training policy, the fidelity of its local risk signal, and the downstream value of using that signal to allocate a fixed budget.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.