acceptodds
Under review as a conference paper at ICLR 2027

Same Delta, Different Effects: Reusing Post-Training Updates after Continual Pretraining

Abstract

Adding a post-training delta after continual pretraining (CPT) can recover instruction following. Does a similar instruction gain mean that the update corrects the same failures at the same cost to language adaptation? We study a fixed delta, defined as the weight difference between a post-trained model and its base, along a Qwen3.5-9B Tibetan CPT trajectory. Overall instruction gains remain similar, while the language-modeling penalty in bits per byte (BPB) grows roughly fourfold. The corrected prompts also change, even within a fixed set that all receiving checkpoints fail. Scaling reveals a related distinction between preserving scores and preserving individual improvements. At roughly half the BPB cost, a scaled update retains 96% of the full update's net IFEval gain but corrects only 81% of the prompts newly corrected by the full update. Its net IFBench gain retention is 31%. Restricting transfer to lower-half feed-forward blocks shows no clear advantage over scaling at matched cost. An external Llama-3.1-8B/Basque comparison also finds instruction gains with a BPB penalty, although that penalty decreases relative to the base point. These results show why successful reuse requires more than recovering an average instruction score: the same update can change which prompts improve and its cost to language-model fit.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.