Tracing the Internal Computation of a Chess Playing Transformer
Abstract
Modern transformer neural networks achieve strong master-level performance in chess, yet a mechanistic understanding of how their internal computations produce move decisions remains limited. Building on advances in mechanistic interpretability, we adapt Low-Rank Sparse Attention (Lorsa) to LC0's policy network and combine it with Transcoders to decompose attention and MLP computations throughout the model into a unified set of sparse features. We also extend circuit tracing to LC0's bilinear policy output head. We introduce a chess-structured taxonomy and an agentic pipeline that proposes feature interpretations and, when possible, translates them into rule-based tests grounded in chess structure. On this foundation, we construct end-to-end attribution graphs that trace attributed pathways from board inputs through sparse features to move logits. Through case studies, we uncover interpretable tactical computations underlying the model's decisions for specific input positions and target moves. Moving beyond position-specific cases, we introduce quantitative metrics showing that computations are largely separated across candidate moves and highly branched within each move, while feature attribution gradually concentrates on the corresponding source and target squares across layers. Our code is available at https://anonymous.4open.science/r/Leela-SAEs-446A.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.