acceptodds
Under review as a conference paper at ICLR 2027

Fighting the Drift: Probe Transfer Across Post-Training Stages

Abstract

Linear activation probes are used for monitoring large language models. Updating a probe for a new checkpoint requires re-extracting activations and may require new labels. First, we study probe transfer across SFT, DPO, and RL in OLMo-3 7B and across SFT and GRPO in Qwen3-4B. On the most disruptive SFT transfers, frozen probes achieve only about half the true positive rate of a probe refit on the target checkpoint, at 1% FPR. DPO and RL updates leave much less headroom for refitting than SFT updates. A refit on the target checkpoint can outperform the probe on its source checkpoint, so we measure transfer performance against the target refit. Label-free representation alignment can close much of this headroom but can plateau below the refit or perform worse than the frozen probe. Second, we develop an orthogonal drift correction that adjusts the old probe, with or without representation alignment, using a small number of labels. With a strong starting probe, correction achieves the best performance at most tested labeling budgets, can match the full refit, and can outperform a from-scratch refit trained on the same labeled data. Cheap probe transfer is a step toward overseeing systems that update faster than their monitors can keep up.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.