From Identifiability to Sample Complexity for Discrete Latent Variables Seen Through Known Channels
Abstract
Causal representation learning seeks to recover meaningful hidden variables from observations. When observations merge hidden states, which properties remain recoverable, and how much data is needed to estimate them? We study this question in finite discrete models. A matrix built from the observation process gives an exact test for whether an average hidden property can be recovered. For every recoverable target, we construct an estimator whose mean squared error has an explicit leading term proportional to the inverse of sample size, plus a smaller correction. A matching lower bound shows that no method improves this leading error in the worst case, and the estimator has a normal large sample limit with optimal variance. When the observation process is unknown but belongs to a finite set, indistinguishable processes with different target values make consistent estimation impossible. With sufficient separation, an adaptive estimator matches the same leading error as if the process were known. Under a common support condition, candidate processes whose observed laws are close, at the scale of one over the square root of the sample size, but whose targets differ leave an error that does not vanish. Simulations and a semi-synthetic census study with a randomized-response channel illustrate these predictions.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.