PhysSAE: Mechanistic Interpretability of Physics-Informed Neural Networks via Sparse Autoencoders
Abstract
Physics-Informed Neural Networks (PINNs) embed PDE residuals into neural network training, but their internal representations remain opaque: it is unknown what physical features their hidden layers encode or whether those features have a localized causal role. We present PHYSSAE, a mechanistic interpretability frame- work that trains overcomplete sparse autoencoders (SAEs) on PINN penultimate- layer activations and evaluates dictionary atoms through direct causal intervention in the original frozen hidden state: , bypassing the SAE decoder entirely. Across six PDE families—heat diffusion, viscous Burgers, Allen–Cahn reaction-diffusion, linear convection (easy and hard regimes), and the nonlinear Schr\"odinger equation—with 3 PINN seeds and 3 SAE seeds each, we show that (i) SAE atoms align with independently-defined physical observables (max Pearson , always permutation null), (ii) the causal footprint of top- aligned atom ablation is 1.2-4.2 more spatially concentrated than PCA or ICA interven- tions, and (iii) top-aligned atoms outperform matched random controls on causal localization for structured physical concepts (ESF80 advantage 0.04–0.44). Two- atom bilateral representations improve concept regression by -0.15 over single atoms, while random pairs decrease it by up to 0.60. These results demonstrate that PINNs develop sparse, physically structured latent representations that can be identified and causally interrogated post-hoc, opening a path toward interpretability-aware scientific machine learning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.