Hierarchical geometry and reference-derived steering in Gemma-2-9B
Abstract
Language models can represent hierarchical relationships geometrically, but the presence of this structure does not establish when it influences an answer. We examine this distinction in Gemma-2-9B across 46 noun hierarchies containing 690 concepts, tracing representations through all 42 transformer blocks. Hierarchical distance structure strengthens in early blocks, while reference-fitted principal components capture broad category distinctions more strongly than finer distinctions. We then add reference-derived category directions to individual concept-token states, excluding query leaves from their own reference fits and using the known target category to orient each edit. At B10, these interventions improve coarse and fine answer accuracy by 2.85 and 3.40 percentage points, respectively. Benefits diminish at later sites despite substantial retained category structure. In a paired cathedral example, a B34 edit strengthens the correct geometric preference while leaving the answer incorrect. State transfers identify the final prompt position as carrying a consequential difference from the successful early edit. Component interventions reveal contributions through this state from later attention and MLP computations, including B23 attention. These findings distinguish hierarchical representation from behavioral influence and show why evaluating geometric interventions requires tracing their effects through subsequent computation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.