Sky sphere representation in language models
Abstract
We analyze whether language models have an internal representation of the night sky that they use in answering questions about proximity on the sky. We find such a representation (a "sky sphere") in residual streams of tested LLMs (size 32B-235B parameters) on prompts like "in the night sky [object X] appears near". In 6 out of 7 models the sky sphere is linearly decodable ( score of 80–90% and median angular error down to – in leave-one-out test). The sky sphere is indeed spherical and not a flat atlas chart, and spectral embeddings of graphs derived from constellation grouping and adjacency or from corpus co-occurrence failed to fit the observed geometry as well as the straight sphere. However the neural sky sphere is bent: including deg-2 spherical harmonics in the model was crucial for the following intervention experiments that tested causal use of the sky sphere. Ablating the 8 directions utilized by the bent sky sphere raises the angular distance to a neighbor provided by LLM's output from - to - in most models (for control, ablating the complementary to the "sky sphere" in top 16 PCA raises the separation by less than ). In steering, when asked for a neighbor of constellation A, replacing "sky" directions with the model fit for a distant constellation B shifts its prediction (probability 60%-85% near A) to the target (47%-83% near B, 0%-4% still near A), which shows active use of the "sky sphere". To our knowledge, this representation is the first example of a (curved) multi-dimensional irreducible feature manifold. Code is released.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.