Hierarchical Identifiability in Multilayer Representation Learning
Abstract
Despite the hierarchical structure exhibited by many real-world generative processes, identifiability theory in representation learning has primarily focused on flat generative models, leaving intermediate representations in hierarchical models largely unexplored. We show that hierarchical structure fundamentally changes the identifiability problem: intermediate representations are governed by admissible inter-layer reparameterizations, which must be controlled across successive layers. We study a class of hierarchical stochastic models with additive nonlinear feature maps, overdetermined linear mixing, and layer-wise innovations. Under explicit nondegeneracy and innovation assumptions, we establish hierarchical identifiability from observational data alone: nonlinear ambiguities at intermediate layers reduce to permutation and component-wise affine transformations, while the root representation is identifiable up to permutation and component-wise local diffeomorphisms. We further show that identifiability can propagate even when only a sufficiently informative subset of intermediate coordinates is affinely identifiable. The proposed structural constraints can be incorporated into simple reconstruction-based estimation. Experiments on synthetic hierarchical data support the qualitative predictions of the theory, while experiments on high-dimensional image data suggest that the resulting identifiability-inspired architectural constraints can provide useful inductive biases beyond the idealized theoretical setting.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.