acceptodds
Under review as a conference paper at ICLR 2027

A Learned Domain-Discriminative Subspace of the Residual Stream

Abstract

Do trained language models represent domain identity, i.e., how the text they are processing departs from their default distribution, in structured form, and is that structure learned rather than an artifact of how we probe for it? We define , a per-layer subspace of the residual stream obtained by joint diagonalization of per-domain generalized-eigenvalue discriminants against generic text, and recover the same compact object in five model families spanning three token-mixing mechanisms. It occupies -% of the hidden dimension, -% in the mixture-of-experts read one sub-layer earlier, sized by the number of calibrated domains rather than by stream width. Because such a construction returns a basis from anything, the depth claim is made at fixed rank and against a floor drawn from its own candidates - an arbitrary subspace of equal rank, which has a rising gate of its own. Against that floor, discriminability along rises with depth in every trained checkpoint, decays in random-initialization controls, and clears it by a near-constant - in all five families. Pseudo-concepts stay an order of magnitude below at every depth and rank. Removing degrades even the least affected calibration corpus - more than the generic reference, whereas a matched-rank variance-only subspace damages domain and generic text indiscriminately. Yet nothing we can measure reads the code, and its causal necessity does not extend beyond the corpora it was calibrated on, even though its discriminative format does. What survives is a learned, compact, corpus-selective code.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.