acceptodds
Under review as a conference paper at ICLR 2027

Brief Treatments Select Long-Run Function in Post-LN Models

Abstract

Can a 25-update treatment change a model's function after thousands of unmodified updates? In an unstable Post-LN decoder trained without warmup, we modify only updates 4–28 of paired runs sharing an exact parent and then resume native AdamW through update 5000. On six complete master seeds, a scaled ascending learning-rate ramp is healthy on 6/6, while an unscaled ascending ramp and a norm-preserving directional intervention are each healthy on 5/6 and fail on different seeds; untreated ACTUAL, an exact-dose constant profile, and a release-continuous V0 profile are healthy on 0/6. Exact-dose reverse and half-LR controls also fail on the original seeds. Together with a successful release-shock control, these comparisons show that the tested outcomes separate strongly by within-window temporal organization beyond cumulative dose, release continuity, and sustained release-step magnitude. Transfer is recipe-dependent: at eight layers, 25-, 50-, and 100-update ramps are healthy on 0/3, 2/3, and 3/3 seeds; at 107M parameters, tested 25-, 100-, and 200-update ramps are 0/3, while Warmup500 remains more robust. At two lower-LR release states, parameter state and complete AdamW state each suffice for healthy continuation, whereas isolated moments do not. Release diagnostics show distinct organizations for healthy directional and scalar routes. These are tested-point results, not a universal warmup law or an early-commitment period.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.