Online Sparse Additive Learning without a Prespecified Reporting Horizon
Abstract
Estimating sparse nonlinear effects from a growing data stream requires updating both the selected variables and their uncertainty. Correlated covariates can give an irrelevant variable a larger initial effect than a relevant one, causing thresholding to fail. The sample means used to center effects also vary across samples. In this article, we propose an online sparse additive method in which selection, estimation, and centering share one kernel-weighted least-squares criterion computed from stored sums. Joint re-estimation with smoothly clipped absolute deviation weights consistently selects the relevant variables in constructed correlated models where initial thresholds fail. Under population selection margins and identifiable correlated effects, continuous-effect estimates attain the known-support optimal rate for a stationary repeat-or-redraw covariate process. With Gaussian autoregressive errors and bounded pairwise density ratios, we establish asymptotically valid joint confidence intervals for a fixed finite set of effect values at each prespecified reporting time. The final sample size need not be specified. For a fixed number of covariates, storage has an asymptotically sublinear bound. In simulations with strong signals, recomputed centers correct the undercoverage of fixed centers. Our method also lowers estimation error against the compared planned-epoch and stochastic-gradient estimators. Bike-rental and appliance energy analyses identify stable sparse nonlinear effects.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.