PLOT: Progressive Localization via Optimal Transport in Neural Causal Abstraction
Abstract
Causal abstraction offers a principled framework for mechanistic interpretability, aligning a high-level causal model with low-level neural computation through interchange intervention analysis. Finding such an alignment, however, often requires fitting and evaluating separate learned mappings across many candidate neural locations. We introduce PLOT, a gradient-free approach for joint correspondence discovery via a global matching of intervention effects. PLOT represents abstract and neural interventions by geometric signatures of their effects on the shared task output and fits an optimal transport coupling between the two collections. The coupling can be calibrated directly into an executable intervention handle or used progressively to select coarse parent sites for finer localization, reducing the cost of searching broad collections of neural sites. Experiments on hierarchical equality, binary addition, and multiple-choice question answering (MCQA) demonstrate fast and accurate direct handles without learning intervention rotations and show that progressive localization can improve accuracy under intervention-size constraints while reducing computational cost. PLOT can also guide gradient-based subspace learning, with PLOT-guided distributed alignment search (DAS) attaining accuracy comparable to full DAS at approximately and lower serial runtime on binary addition and MCQA, respectively. PLOT thus separates the search for candidate correspondences from optional subspace learning and provides a common framework for constructing and testing neural intervention handles for a given causal model.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.