acceptodds
Under review as a conference paper at ICLR 2027

Spectrally Broad, Behaviorally Compressible: Functional Cores of Safety-Tuning Updates and What Downstream Fine-Tuning Does to Them

Abstract

Safety tuning reliably reduces large language models harmfulness, but the geometry of the parameter change it induces, and how that geometry relates to the loss of safety under later fine-tuning, remain only partly understood. We investigate the underlying geometric structure of safety tuning and examine how downstream fine-tuning erodes alignment, both through a single object: the functional core of the safety update. We show that the safety update is spectrally broad but behaviorally compressible. Moreover, we find that this functional concentration appears early and persists even as the underlying subspace continues to evolve. We then follow the core through downstream fine-tuning in the same safety-derived coordinates. Its small-energy in-core displacement opposes the learned safety direction, and signed interventions establish the behavioral relevance of these coordinates. We then use the discovered core as a correction space: a current safety-memory loss guides minimal adjustments to downstream updates. This improves endpoint safety while largely retaining the downstream utility gain. Together, these findings connect the acquisition, erosion, and correction of safety through a learned functional core.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.