Interpretable Deep State Space Models via Semantic Alignment
Abstract
Deep state-space models (DSSMs) are promising tools to control complex dynamical systems, but their latent states are typically unintelligible to humans and therefore difficult to inspect, steer, and debug. In response, we propose a general methodology for constructing expressive and intrinsically interpretable DSSMs that permit the alignment of both their state representations and dynamics with human semantics. Specifically, we decompose a DSSM's internal representation into an expert-defined dynamical concept state and a residual latent state, which is learned to compensate for the predictive information that is missing from the concepts. Notably, the resulting formulation supports interventions that propagate through the learned dynamics and steer the subsequent concept trajectory. We instantiate the methodology in both deterministic and stochastic model variants and evaluate them across four dynamical systems, demonstrating free-running concept forecasting accuracy on par with opaque DSSMs while also enabling interventions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.