Verbalizing Multi-Token Concepts in LLMs
Abstract
Lens methods inspect model computation by reading hidden states as ranked vocabulary tokens. Yet the concepts humans need to read out often span multiple tokens—entities, phrases, intermediate objects—making token-level readouts incomplete. Reliable multi-token readout with little model-specific preparation remains challenging. We introduce Concept Lens: token-level lens clues guide candidate search, while the model recovers a representation for each candidate and scores it against the source activation in the same readout space. Across 2,400 multi-hop clozes on five LLMs (8B–70B), Concept Lens instantiated with J-lens and R-lens achieves average Rank@10 scores of 36.6% and 54.5%, respectively, compared with 21.7% for Template Lens. Controlled tests show that recovered representations preserve concept identity and that concept rankings track changes in the source activation. Concept-swap interventions further shift model answers toward those associated with the replacement concepts. Our results move lens-based interpretability beyond token-level inspection toward the concept-level structure of model computation. Our code is available at https://anonymous.4open.science/r/ConceptLens.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.