acceptodds
Under review as a conference paper at ICLR 2027

Nine Emotion Anchors Recover a Valence Direction: No-Label Cross-Modal Transport and Scale-Aware Steering

Abstract

Inside a modern language model sits an internal direction that tracks how positive or negative a sentence feels. We recover this valence axis (V-axis) from nine generic emotion-category anchors – no downstream sentiment labels – and follow it through four questions: can we recover it cheaply, is it broadly readable, does readable imply causally writable, and where does it fail? Recover: the recipe averages ∼50 emotion-anchored paragraphs per category into nine centroids and takes their top principal component (∼ 384× fewer labels than a supervised SST-2 probe: n=18 supervision events vs. 6,920). At matched budget the nine anchors – not the specific reader – carry the signal: difference-of-means, CAV, and LDA readers of the same anchors match or beat the principal-component recipe, all clearing a 500-direction random null on 8/9 models. Read: the axis is stable when explicit emotion words are stripped from the prompts (cosine 0.98 between original and word-removed axes), reads four disjoint sentiment corpora above null (SST-2, IMDB, Rotten-Tomatoes, Yelp), and appears in a self-supervised vision encoder that never saw language (DINOv2, valence AUC 0.790 vs. CLIP control 0.972). Write: readable does not mean freely writable. Removing the single axis leaves a retrained probe intact – valence spans a 4–36-dimensional redundant subspace – yet the model still causally uses it at inference time: directional ablation (Arditi et al., 2024) of the axis drops the model’s own sentiment output on 6/9 families while a norm-matched random-direction ablation moves it ≈0. Steerability follows a residual-stream-norm regularity that predicts held-out models (leave-one-model- out R2=0.88), and an explicitly-defined donor→receiver map transports the axis across models above a random-map null (3 model pairs, all 6 off-diagonal cells, 20–31×). Bound: the recipe is at-or-near chance on categorical concepts (AxBench Concept-500 (Wu et al., 2025), KS p=0.21) and on arousal, and on EEG valence is decodable but its geometry is not shared with machine encoders – a boundary case, not a fourth shared modality; a geometric “admissibility” criterion we tested does not predict held-out success and is reported only as a negative. Valence is thus the rare high-level attribute recoverable from category anchors alone that also transports without target labels across four independently-trained encoders and stays causally writable under a scale-aware fix – a reach that is real but, we show, bounded by residual-stream scale and cross-encoder geometry.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.