acceptodds
Under review as a conference paper at ICLR 2027

Shared Parametric Geometry and Data Provenance in Model Organisms for Emergent Misalignment

Abstract

Recent work has discovered that large language models (LLMs) can develop broadly misaligned behaviors after being fine-tuned on narrowly harmful datasets, a phenomenon known as emergent misalignment (EM). However, the fundamental mechanisms enabling such harmful generalization across disparate domains remain poorly understood. We therefore adopt a geometric perspective to study EM in the weight space, starting with established LoRA-based model organisms that exhibit a cross-task linear structure: different, narrowly harmful supervised fine-tuning tasks converge to shared low-dimensional parameter subspaces and are behaviorally equivalent via linear mode connectivity, with interpolations between the resulting models maintaining coherent, broadly misaligned behavior. We further test the robustness of this geometric structure across LoRA ranks and model sizes, while also examining benign controls, data-generation templates, and data-generating LLMs, revealing a key role for data provenance: harmful and benign fine-tunes sharing data-generation templates exhibit comparable subspace overlap, whereas the tested cross-template comparisons yield little overlap. We subsequently examine the implications for safety through two post-hoc vaccination strategies, negation and orthonormal projection, finding that both improve alignment across the tested model families and that EM-derived projections yield greater alignment gains on average than benign-derived projections, while orthonormal projection also better preserves reasoning performance than negation. Collectively, our analyses connect data provenance to shared parametric geometry, showing that geometric convergence need not be specific to misalignment but can instead reflect common structure inherited from the data-generation process, while remaining behaviorally consequential for targeted interventions in LLM safety.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.