acceptodds
Under review as a conference paper at ICLR 2027

Localizing Behaviors from Generations, Not MCQA Options

Abstract

Task settings that elicit a single-token response from a Language Model are widely used in interpretability research to causally localize the circuits underlying the task. Yet, this formulation is poorly suited to complex behaviors such as factual recall, reasoning, summarization, and style transfer. Prior work therefore often studies these behaviors through multiple-choice question answering (MCQA) formulations, or leaves them unexplored altogether. In these constrained settings, a model is said to exhibit the target behavior when it predicts the aligned answer choice. We propose, and defend, an alternative: localizing behavioral circuits directly from free-form generative responses. Across five model families and four tasks with matched MCQA and generative variants, we find that circuits localized from free-form responses are more precise than their MCQA-derived counterparts. We employ mechanistic analyses to explain this difference: MCQA introduces intermediate computations that are absent during direct generation, and contains high-rank representations. Taken together, our results suggest that localizing circuits from free-form responses provides a more direct and effective basis for model control.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.