acceptodds
Under review as a conference paper at ICLR 2027

Cross-Layer Transcoders as Hypothesis Generators for Causal Circuit Analysis

Abstract

Cross-layer transcoders (CLTs) replace transformer MLP computation with sparse, interpretable feature dictionaries that support prompt-specific attribution graphs. Their graph scores, however, are derived from a local linearization and need not predict the effect of intervening on a feature in the full nonlinear model. We study this gap in \model. We train a -feature latent-mixing CLT that reconstructs all MLP outputs while using fewer decoder parameters than a direct cross-layer parameterization, then construct direct-effect and Neumann multi-hop attribution graphs for factual and mathematical prompts. Across five held-out evaluation distributions, reconstruction MSE ranges from to . The graphs surface semantically coherent recurrent candidates, but their Neumann scores do not directly recover measured single-feature effects. Direct perturbations identify \feature as a robust sentence-boundary continuation feature rather than a feature specific to However. During AIME generation, amplifying this feature reduces the frequency of newline continuations after periods from to and increases a restricted family of connective continuations led by Since and We. This local shift changes global solution form: across matched generations, mean lines per solution fall from to and mean line length rises from to tokens. Accuracy changes are modest in aggregate and heterogeneous by problem; amplification raises AIME 2024 Problem 12 accuracy from to but is not uniformly beneficial. These results support CLT attribution graphs as tools for candidate discovery while showing that causal claims require explicit intervention.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.