Cross-Layer Transcoders as Hypothesis Generators for Causal Circuit Analysis
Abstract
Cross-layer transcoders (CLTs) replace transformer MLP computation with sparse, interpretable feature dictionaries that support prompt-specific attribution graphs. Their graph scores, however, are derived from a local linearization and need not predict the effect of intervening on a feature in the full nonlinear model. We study this gap in \model. We train a -feature latent-mixing CLT that reconstructs all MLP outputs while using fewer decoder parameters than a direct cross-layer parameterization, then construct direct-effect and Neumann multi-hop attribution graphs for factual and mathematical prompts. Across five held-out evaluation distributions, reconstruction MSE ranges from to . The graphs surface semantically coherent recurrent candidates, but their Neumann scores do not directly recover measured single-feature effects. Direct perturbations identify \feature as a robust sentence-boundary continuation feature rather than a feature specific to However. During AIME generation, amplifying this feature reduces the frequency of newline continuations after periods from to and increases a restricted family of connective continuations led by Since and We. This local shift changes global solution form: across matched generations, mean lines per solution fall from to and mean line length rises from to tokens. Accuracy changes are modest in aggregate and heterogeneous by problem; amplification raises AIME 2024 Problem 12 accuracy from to but is not uniformly beneficial. These results support CLT attribution graphs as tools for candidate discovery while showing that causal claims require explicit intervention.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.