FIRE: Learning Identifiable Operators for Multi-Environment Generalization
Abstract
Multi-environment data mix shared subject structure with environment-specific distortion. Domain generalization targets robust prediction and harmonization removes environment effects; neither is designed to identify the shared latent directions. Factored Identifiable Representation via Environments (FIRE) models each environment's observations as unknown full-rank linear transforms of a common subject representation. Stage 1 decomposes a reference environment's third cumulant; Stage 2 freezes the corresponding recovered rank-one operators during supervised prediction. Independent loadings with nonzero third cumulants and noise with zero third cumulant identify directions up to sign, permutation, and reference coordinates from one environment, without pairing. Under sub-Gaussian tails, a whitening-free estimator attains error up to a conditioning factor once ; without structural loading constraints, identification is impossible. Paired cross-environment second moments cancel independent noise. Synthetic recovery follows scaling, tolerates subject-dependent maps and poor conditioning, and degrades under nonlinear or spatially heterogeneous maps. Updating operators raises recovery error from to ; freezing preserves geometry without increasing prediction loss. Against ERM and six domain-generalization methods using identical losses, features, and splits, FIRE has the highest mean out-of-distribution accuracy on four of five benchmarks, with gains modest relative to seed variability. Random frozen directions predict equally well on Camelyon17. Spectral recovery supplies identifiable structure: its TCGA-BRCA directions separate PAM50 subtypes more strongly than random or PCA directions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.