Hidden States as States: Discretizing Hidden-State Geometry in Language Model Computation
Abstract
While large language models (LLMs) operate on discrete tokens, their internal computation unfolds in a continuous, high-dimensional representation space. We introduce Hidden States as States (HSS), a training free and label free framework that discretizes hidden representations at a chosen granularity into a cross-layer vocabulary of geometric states. HSS clusters representations within each layer and aligns cluster centroids across adjacent layers, converting each forward pass into a layer-indexed discrete state sequence. The resulting maps are stable under clustering perturbations, expanding in intermediate layers and contracting later in model-specific patterns. Conditioning on behavioral labels further reveals distinct state occupancy patterns, with failures collapsing into a few persistent sink states that emerge in early layers. Simple count-based predictors over these discrete state sequences outperform closed-form continuous baselines, match trained continuous probes on capability and safety benchmarks, and support online failure warning during generation. Reconstruction and activation editing show that, despite their coarse granularity, these states retain functionally relevant structure for downstream computation. HSS thus converts continuous hidden dynamics into an actionable discrete structure.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.