acceptodds
Under review as a conference paper at ICLR 2027

CAUSAL ABSTRACTION OF MAGNITUDE IN LAN- GUAGE MODELS

Abstract

Language models can distinguish “5 degrees Celsius” from “35 degrees Celsius,” but whether such quantitative concepts are reusable latent variables, rather than context-specific surface features, is unclear. We test this via cross-form causal interchangeability (transplanting a magnitude direction across linguistic surface forms) and multi-consequence coherence (whether one intervention coherently shifts many independent consequences, not just one), against random, output-specific, and novel-context controls. On Qwen2.5-7B-Instruct, temperature is strongly decodable (), and a magnitude direction transfers almost as well across forms as within one: mean coherence over 15 consequences is same-form and cross-form, against a random-direction null of – a gap of about pooled standard deviations, confirmed by a non-overlapping bootstrap 95% CI, and equally strong for a novel context category and a continuous dose-response. Reaching this required catching two of our own bugs: an all-position intervention let a random direction mimic the real effect, and 3 of 15 consequences were authored with the wrong sign, caught only by the dose-response check itself. Extending to a second concept and two more models gives a mixed picture: distance replicates far more weakly (1–2 pooled SD), Phi-3.5-mini-instruct shows a moderate effect (2 SD), and Mistral-7B-Instruct-v0.3 shows none (0 SD) despite decoding temperature more accurately than Qwen. The central claim rests on this evidence chain, not on an output-specific-steering control, which stayed unreliable after two fix attempts; a small held-out set and an unresolved context-magnitude compositionality test are the main remaining gaps. Probing alone can overstate conceptual representations: decodability is not equivalent to causal reusability, and causal abstraction – cross-form transfer plus multi-consequence coherence – is a stronger criterion for studying whether representations support flexible computation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.