acceptodds
Under review as a conference paper at ICLR 2027

Manifestation Units: Representing Mechanistic Findings for Access and Intervention

Abstract

Mechanistic interpretability produces heterogeneous findings about model components, relationships, and causal effects, but these findings are often left in analysis-specific forms that complicate downstream reuse. We introduce **Manifestation Units**, a persistent structured interface that organizes findings through shared semantic roles while retaining architecture-specific evidence. We evaluate MUs through Structural Accessibility, whether recorded findings can be recovered under bounded access, and Causal Utility, whether represented numerical information preserves downstream component selection and intervention effectiveness. Across VAE, CNN, and GPT-2 settings, the MU pipeline achieves the highest answer-slot accuracy among nine matched BM25 conditions (54.58%, 79.44%, and 79.72%). In the CNN, MU-mediated access reaches 92.37% targeted success, while complete-access MU and native findings produce identical selections and activation-patching effects. In GPT-2, MU preserves the complete native head ranking and seven-head selection, which achieves 61.87% normalized recovery on fresh IOI examples versus 6.30% for layer-matched random controls. These results support treating persistent representation between mechanistic analysis and downstream use as an explicit and independently evaluable layer.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.