acceptodds
Under review as a conference paper at ICLR 2027

Geometric Alignment: Training Against Concentrated Behavioral Geometry

Abstract

Many alignment interventions assume undesirable behavior lives in a compact representation direction. Before removing it, we ask whether a coherent direction exists and whether ordinary training can avoid it without an inference-time edit. We answer with three falsifiable gates—concentration (A), geometry-held-out exclusion after geometric training (B), and an opener-CE diagnostic (C)—and the training bridge L = L_task + λ G_bad, exact for linear readouts as w^T Σ_bad w. On a planted residual-stream world the chain passes under geometric training with no post-hoc scrub (held-out response drop 0.989). On a Gemma-4-12B educational proxy, KL-regularized LoRA with a projected-energy penalty cuts bad-subspace energy by 32.5% on a six-prompt report slice while the specified opener-CE diagnostic stays within its Gate C band, again with no activation scrub; bridge CE stays near-flat, so Gate B is energy, not CE. The same A→B→C design on Qwen2.5-7B-Instruct (same prompts and splits) yields a 31.8% report-slice energy drop and opener-CE +0.1%: a protocol replication, not evidence of universal generalization. A second family (sycophancy, Gemma) is NO-GO at layer 16 and GO (−44.3%) at layer 48: concentration licenses the question, but layer still matters. On these prompt families, concentrated displacement geometry can be suppressed in projected energy without moving that opener-CE diagnostic out of band; if concentration fails, the gates say NO-GO.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.