BICD: Bidirectional Circuit Discovery by Dynamic Forward Selection and Backward Pruning
Abstract
Mechanistic interpretability (MI) has developed many mature methods for discovering subcircuits in large language models. However, methods that abstract an LLM as a computational graph often face an enormous number of edges and consequently inefficient pruning. Methods based on activation patching, on the other hand, depend heavily on manually specified patching interventions, and the subcircuits they discover often have limited ability to perform the task independently. Drawing on the complementary ideas of forward search and backward pruning, we propose Bidirectional Circuit Discovery (BiCD), a simple and efficient bidirectional search algorithm. BiCD seeks to obtain a subcircuit that is as small as possible while preserving as much of the model's functionality as possible.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.