Spectral Deconfounding for Causal Inference from High-dimensional Perturbed Model
Abstract
High-dimensional causal inference from observational studies is challenging due to unobserved confounders. To ensure identifiability, previous works assume the dense confounding condition, which means that each latent confounder is correlated with most covariates. This dense confounding introduces strong correlations among covariates, leading to large singular values along certain directions of the design matrix. To remove the confounding effect, spectral deconfounding methods apply spectral transformations to the design matrix, attenuating the dominant singular components caused by latent confounders. In sparse linear regression, it has been established that such deconfounding can improve the estimation for Lasso. However, a fundamental question remains open: Can spectral deconfounding identify which features are causally relevant to the response? This is known as the model selection problem and is important in many applications like scientific discovery. In this paper, we show that by weakening confounding-induced spurious correlations, spectral deconfounding enables the irrepresentable condition holds asymptotically, therefore enabling Lasso to identify and optimize within the causal features set. This perspective in turn motivates us to reduce the bias produced by the penalty, and thus improving the estimation. Specifically, we employ the Linearized Bregman Iteration, which was proposed as an optimization algorithm and was recently found to enjoy the unbiased property from a differential-inclusion perspective in the fixed design setting. We extend this property to random design setting with latent confounders, and show that, our estimator achieves the oracle estimation rate under the beta-min condition, matching the minimax-optimal rate for unconfounded setting with fixed design up to a logarithmic factor.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.