acceptodds
Under review as a conference paper at ICLR 2027

Tracing Low-Level Linguistic Properties Through Sparse Representation to Behavior

Abstract

Language models (LMs) are built on human languages; therefore, the structural properties of languages are deeply embedded in a model's representations, affecting output-level behaviors. The relationship between these effects (as well as their nature) and the linguistic information remains largely unclear. We study this relationship through a combination of approaches across 40 low-level linguistic variables represented by 20,000 controlled contrast pairs across eight languages. In XGLM-564M, we compare supervised linear probes with features derived from a BatchTopK sparse autoencoder trained independently on multilingual natural text. Through evaluating these features across a set of metrics, performing ablation, and steering, which is then followed by next-token-level behavior evaluation, we find a set of six variables that show significant and consistent results across most tests while major subsets survive across different assessments. In this study, we aim to highlight the possibilities behind the broader intersection of computational linguistics and mechanistic interpretability, while showcasing how not just the features themselves, but the methods and approaches we use also produce significantly overlapping evidence and localization of representations.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.