BEYOND FIRST CROSSINGS: PRESERVATION IN FACTUAL REWRITING AT MATCHED MEAN OPTIMIZATION-PROMPT LOSS
Abstract
We study which preservation differences in factual rewriting persist after match- ing the mean negative log-likelihood (NLL) achieved on optimization prompts. Starting from the same complete intermediate state, we apply suffix programs with different learning rates and compare discrete first crossings with directly evaluated parameter states within the final update. An initial exposure-matched load–cost reversal is absent under matched base-training steps, motivating these within-state comparisons. In five new controlled initialization–mapping pairs, subsequent mean-loss matching retains high-load old net-cost reductions of ap- proximately 1.15 and 1.07 nats for slow relative to fast suffixes across two pre- fixes. Here, old net cost is the change in mean old-target NLL from the base model. On Qwen3-8B-Base/CounterFact, mean coarse-step contrast magnitudes shrink from roughly 0.03–0.06 nats to approximately 0.0063 nats or less, whereas fine-step comparisons retain much smaller differences of approximately 0.002– 0.004 nats after extending eligible trajectories and narrowing the matching band. In the fine-step narrow-band comparison, slow suffixes have lower mean old net cost but higher mean held-out test NLL on complete-expression batches, despite equal test-accuracy counts in every paired batch. Thus, the large controlled dif- ferences survive mean-loss matching, whereas the coarse natural-fact contrasts largely do not. Per-edit losses and independent-expression effectiveness can still differ at a matched mean optimization loss.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.