acceptodds
Under review as a conference paper at ICLR 2027

Interpretable Deep State Space Models via Semantic Alignment

Abstract

Deep state-space models (DSSMs) are promising tools to control complex dynamical systems, but their latent states are typically unintelligible to humans and therefore difficult to inspect, steer, and debug. In response, we propose a general methodology for constructing expressive and intrinsically interpretable DSSMs that permit the alignment of both their state representations and dynamics with human semantics. Specifically, we decompose a DSSM's internal representation into an expert-defined dynamical concept state and a residual latent state, which is learned to compensate for the predictive information that is missing from the concepts. Notably, the resulting formulation supports interventions that propagate through the learned dynamics and steer the subsequent concept trajectory. We instantiate the methodology in both deterministic and stochastic model variants and evaluate them across four dynamical systems, demonstrating free-running concept forecasting accuracy on par with opaque DSSMs while also enabling interventions.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.