acceptodds
Under review as a conference paper at ICLR 2027

LaTense: Measuring the Limits of Geometric Activation Steering

Abstract

Claims that activation steering causes or prevents "latent collapse" are routinely supported by n-gram repetition rates. We show that this metric is governed by output length rather than by the intervention. Across 19,268 generations from 56 runs spanning six models from 0.5B to 9.2B parameters, three families, three benchmarks, and five steering configurations, generated length and token 3-gram repetition correlate at Spearman rho = 0.881 (Pearson r = 0.907): mean repetition rises monotonically from 0.00% for generations under eight tokens to 49.88% above 549 tokens, while the gap between steered and unsteered decoding never exceeds 1.1 points and a 15x increase in parameters moves it by under 2.5. A correct polynomial long division scores 68.9% repetition. A shuffled-token control shows why: the metric conflates task-intrinsic structural repetition, such as the restated scaffolding of a long derivation, with degenerate looping, and only the latter is a failure. We reach this through a controlled study of LaTense (Latent Sense), a dynamic steering framework that modulates intervention strength per token from the local geometric alignment between the hidden state and a reasoning vector, evaluated on Llama 3.1 8B Instruct, Gemma 2 9B IT, and Qwen-2.5-7B-Instruct against unsteered decoding and static Contrastive Activation Addition (CAA) at matched intervention layers and matched decoding. The method helps on one model-task pair of nine (Gemma 2 9B IT on StrategyQA, 74.20%, +8.60 over static CAA, p < 0.0001, at 122.7 generated tokens per problem) and is neutral or harmful elsewhere; its cosine gate is inert in practice, since measured alignment stays near cos(h,v)   0.03, and a norm-scaling-only ablation reproduces that single gain exactly. We conclude that a single linear direction confers limited benefit regardless of how it is modulated, and that repetition-based collapse metrics must report output length, or compare at matched length, in order to be interpretable. Code and evaluation scripts are available in the supplementary material and at https://anonymous.4open.science/r/latense.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.