SFT Leaves a Reproducible Signature in Activation Geometry: Calibrating Post-Training Interpretability Against Run-to-Run Noise
Abstract
A model diff credits post-training with whatever separates a base checkpoint from one post-trained descendant, presuming a rerun would land close by. On the sparse co-activation graph of Qwen3-1.7B it does not: two LoRA-SFT reruns sit over half as far apart as either sits from the base. We score each base-to-post contrast against its own recipe’s run-to-run band, and show that under run exchangeability k runs certify no better than p ≤ 1/(k + 1) however large the effect, so the band’s z is a descriptive scale. SFT leaves a large, reproducible signature: base→SFT displaces the graph by 19.9 pairwise-distance SDs beyond the mean distance between four independent LoRA-SFT runs, and the pre-registered estimator replicates on Llama-3.1-8B with Llama-Scope. At an identical 29 steps, four SFT and four DPO runs occupy disjoint displacement ranges, a gap 5.3× either arm’s spread, without any band. Three other readings stop short of a verdict: DPO’s band membership is dose-dependent (0/4, 2/4, 4/4 contrasts outside at 12, 29, 60 steps) while its displacement stays flat; light GRPO is a secondary non-detection, not an absence, and we run no equivalence test; and the graph concentration read as an RL signature is already in the zero-RL base, and a dictionary-width control removes an apparent scaling law in it. A pre-registered causal sweep (three edge budgets, 4,816 held-out prompts) found the displaced sources load-bearing (12× a matched random control in mean, 3.5× in median at 10,000 edges), but their stage-specificity interaction reversed at two budgets and the third failed its faithfulness gate, so we claim calibrated structural displacement, not a changed mechanism. Model diffs should be read beside their recipe’s rerun band, within-band diffs as non-detections. Scope: the full detection protocol covers one backbone (Qwen3-1.7B), one SFT recipe and general-text probe states; the matched-base control replicates across three families.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.