Manifestation Units: Representing Mechanistic Findings for Access and Intervention
Abstract
Mechanistic interpretability produces heterogeneous findings about model components, relationships, and causal effects, but these findings are often left in analysis-specific forms that complicate downstream reuse. We introduce **Manifestation Units**, a persistent structured interface that organizes findings through shared semantic roles while retaining architecture-specific evidence. We evaluate MUs through Structural Accessibility, whether recorded findings can be recovered under bounded access, and Causal Utility, whether represented numerical information preserves downstream component selection and intervention effectiveness. Across VAE, CNN, and GPT-2 settings, the MU pipeline achieves the highest answer-slot accuracy among nine matched BM25 conditions (54.58%, 79.44%, and 79.72%). In the CNN, MU-mediated access reaches 92.37% targeted success, while complete-access MU and native findings produce identical selections and activation-patching effects. In GPT-2, MU preserves the complete native head ranking and seven-head selection, which achieves 61.87% normalized recovery on fresh IOI examples versus 6.30% for layer-matched random controls. These results support treating persistent representation between mechanistic analysis and downstream use as an explicit and independently evaluable layer.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.