acceptodds
Under review as a conference paper at ICLR 2027

Read–Write Factorization: Preference Geometry for Personalized Decision Making

Abstract

Personalizing large language models requires both *reading* a preference from context and *writing* it into a decision, yet activation steering typically ties these operations. We show that this tying fails along three geometric degrees of freedom: *where* to read and write, *how many* coordinates are needed, and *which directions* support readout versus causal control. Across two 7B architectures, extraction-by-application intervention surfaces are approximately separable, with rank 1 explaining 81–91% of variance in homogeneous-preference conditions. Heterogeneous team preferences are low-rank but not scalar (effective rank 3.14), with the dominant coordinate tracking preference severity (ρ = 0.952). Most strikingly, high-decoding read directions are nearly orthogonal to a gradient-defined write direction (|cos| ≤ 0.026), and their span captures only 0.07% of its squared norm: decodability does not determine causal control. A sign-symmetric decomposition further separates directional control from sign-invariant perturbation. Together, these results establish *read–write factorization*: preference representation and causal control need not share a single steering geometry. Behaviorally, the dissociation transfers to PrefEval's original implicit-preference classification task: write but not read steering improves accuracy (34% → 51%). On a CL-Bench contextual-binding task, separately parameterizing representation and coupling raises binding accuracy from 46.7% to 90.5% ± 1.5%, with ablations attributing most of the gain to coupling the context-dependent representation rather than to the learned direction alone.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.