What the Nonnegative Lens Misses: Energy-Matched Causal Tests of the J-Space in Open Language Models
Abstract
The Jacobian lens maps residual activations onto token-indexed output directions; "J-space" is a sparse nonnegative reconstruction of those atoms. We test how much causal work that reconstruction does on open-weight models with preregistered energy-matched interventions. When residual energy is matched, nonnegative J installs less of a contrastive concept shift than the leftover R=D−J. That R >> J ordering holds on the causal cells that report J−R—six families, rich and legacy reductions (future-summed as robustness), base models, and scales where capability permits ( 7B–32B)—and on a Gemma-4 invented-registry copy the remainder still transfers donor answers while reconstructed-J arms transfer none. Geometric capture stays in a 0.10–0.23 band from 0.5B to 32B. Nesting reconstructions on one dictionary is a selection-partition check: nonnegative OMP matches unconstrained least squares on the same support, while signed OMP raises reconstruction by 1.86–5.24 times at k=100 and flips J versus remainder on geography families. Increasing nonnegative budget fails to close the gap. Installation is not necessity: on Qwen2.5-32B-Instruct (n=28, descriptive), a post-hoc 2×2 of holding×answer-exposure shows an interaction, not an unconfounded holding effect; signed selection is unselective. Held-out known-attribute transport is strong for the full shift (0.82–0.93) and weak for the official nonnegative readout (0.15–0.23). The lens atoms track model knowledge; under the tested nonnegative-OMP reconstruction of contrastive shifts (rich-transport primary; future-summed as robustness), that code discards most of the equal-energy installable component. Preregistrations, frozen manifests, and code are released.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.