acceptodds
Under review as a conference paper at ICLR 2027

Signatures of semantic search in the activations of large language models

Abstract

When recalling lists of concepts (e.g., animals) during the semantic fluency task (SFT), both humans and large language models (LLMs) organise their output into clusters of related items (e.g., sea animals) that are punctuated by strategic switches between clusters. In humans, this pattern can be explained by a semantic foraging process, whereby distinct neural and behavioural signatures accompany within-cluster production (“exploit") and between-cluster switching (“explore"). Whether LLMs likewise represent these two search regimes within their internal states is unknown. Here, we apply a range of mechanistic interpretability techniques to provide evidence for semantic foraging in LLMs. In Study 1, we show that concept-level activations predict switching in SFT. First, we find that switching coincides with elevated next-token uncertainty (low token probabilities). We then show that these switch-associated differences in activations emerge in the early-middle layers by using the recent Jacobian lens (J-lens) approach, which maps intermediate-layer residual-stream representations to token-level activations. Moreover, the probability of switching rises as the set of strongest J-lens activations (the J-space) becomes depleted of items from the category currently being produced, analogous to explore-exploit decision-making during human semantic foraging or animal patch foraging. We then show that middle-layer J-lens activations of abstract category-related labels (e.g., “water") increase in anticipation of switching into that category. We confirm that these intermediate representations causally influence switching by using them to derive steering vectors that target switching into particular categories. In Study 2, we identify generic residual stream directions that are activated during and in anticipation of switching. By steering model activations along anticipatory switching directions, we bias increased or decreased rates of switching. Our paper extends the semantic foraging framework to artificial intelligences and provides evidence that LLMs maintain distinct representational signatures for exploration and exploitation as they verbalise conceptual information.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.