acceptodds
Under review as a conference paper at ICLR 2027

Model Internals Speak Volumes: J-Lens Loudness Predicts Decodability

Abstract

When probing the activations of a Large Language Model (LLM), several design decisions arise. Besides the probe design itself, one must choose which layer(s) and token position(s) to probe. These choices are often made heuristically or through expensive sweeps over candidate layers and positions, and to our knowledge, no method yet predicts the decodability of a signal directly from the activations. We take a step in this direction by introducing J-Lens loudness, a new application of the Jacobian lens (J-Lens) that guides these choices without labels or probe training. We represent a signal of interest by a small subset of the model's vocabulary, and define J-Lens loudness as the log probability mass that the J-Lens assigns to this subset at a given activation. In the reasoning chains of two LLM agents navigating grid worlds, the balanced accuracy of next-action probes increases steadily with direction signal loudness. Taken together, we show that J-Lens loudness can be used as a criterion to find layer–token positions whose activations contain underlying computation relevant to a signal of interest.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.