Evaluating Monitor Reuse After Language-Model Updates: Ranking Targets and Audit Evidence
Abstract
After a language-model update, should a monitor that predicts its behavior be kept or retrained? Our monitors predict whether a reasoning problem and an intended answer-preserving rewrite receive different answers. We compare a training-stage change with a within-stage revision in OLMo-3 7B and Tülu 8B. Three measurements bear on the decision, and each can mislead alone. Absolute quality: pooled AUROC rewards recognizing the benchmark. In Tülu, a benchmark-only predictor reaches pooled AUROC .684 with within-benchmark (macro) AUROC exactly .5 by construction, and a hidden-state probe with macro AUROC .701 retains a pooled–macro gap of .0735 (fixed-score 95% interval [.0296,.1323]). Relative comparisons: on fixed OLMo predictions, the major-minus-revision refitting advantage is −.0099 under pooled and +.0186 under macro AUROC. An exact decomposition attributes the difference to ranking pairs and weights; resampling leaves the population sign pattern unresolved. Audit evidence: with 25 labeled template groups the specified OLMo audit returns no quality bound, so its zero decision error reflects default distrust. The reused Tülu probe retains macro AUROC .689 against .723 after refitting, yet 50-group audits of its five saved fits approve reuse in at most 24% of samples. Removing the audit's sample partitioning returns more bounds and approvals, including false approvals against finite held-out references. We recommend reporting absolute performance under a stated ranking target, and bound availability, together.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.