acceptodds
Under review as a conference paper at ICLR 2027

Structuring Neural Representations with Geometric Supervision

Abstract

Can mechanistic understanding of language models enable training with targeted supervision of internal representations? We leverage previously uncovered neural geometry to design training objectives that sharpen the representations of task-relevant variables. First, we fine-tune Llama-3.1-8B to answer questions such as “What month is four months after June?” by directly supervising representations of the input month (June = 6) and offset (4), which are encoded with Fourier harmonics in the residual stream. Second, we fine-tune the model to predict line breaks in text wrapped to a fixed width by directly supervising the number of characters since the preceding newline, which is represented by a low-dimensional helical structure. In both experiments, we achieve parity with standard fine-tuning using only internal supervision with no output-level task loss. Moreover, the success of geometric supervision can not be isolated to the targeted representations because patching in those representations is not sufficient to recover the performance gain. Together, these findings show how mechanistic understanding of neural geometry can inform training.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.