ECRL: Explainable Causal Representation Learning via Semantic Anchoring of Exogenous
Abstract
Understanding dynamical systems requires not only accurate prediction but also representations that reveal which external mechanism drives each state change, enabling mechanism-level explainability. Yet predictively equivalent causal representations may assign different latent blocks to the same named mechanism, leaving the semantic meaning of block-level interventions unresolved. We study this semantic orientation ambiguity using anchor environments that reveal which exogenous block changed, but not its realized value. Under affine residual ambiguity, known block dimensions, and block-spanning mean shifts, we show that off-target response bounds cross-block mixing, with zero response identifying named blocks up to within-block invertible affine transformations. We derive a finite-sample certificate and extend the result to exact temporal innovations. Motivated by this theory, ECRL localizes anchor responses in transition innovations while sharing coordinates with the state, decoder, and predictor. Known-affine calibration validates the certificate, while full-span neural experiments evaluate factor-value-free intervention-effect prediction. Across five controlled SCM families, 40 seeds, and eight held-out mechanism banks, ECRL reduces effect MSE from 1.1442 to 1.0856 versus a capacity-matched CITRIS-style model, a 5.12% reduction with a two-way cluster-bootstrap interval of [0.0438,0.0740]. ECRL thus links conditional semantic orientation to mechanism-specific effect prediction.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.