acceptodds
Under review as a conference paper at ICLR 2027

Imputers Improve the Wrong Features: Placement, Not Reconstruction Error Alone, Governs What Better Imputation Buys

Abstract

Imputers are compared by reconstruction error, and better reconstruction is assumed to buy better prediction. In practice the gain is small, and this paper measures where the improvement goes. Holding the total reconstruction error fixed and moving it across features changes downstream accuracy substantially, because predictive importance is concentrated in a minority of features, while the standard imputers we evaluate spread their improvement evenly and place only about a third of that improvement where the importance sits. We then attribute the gain using the imputations themselves: replacing only the important third's imputed values recovers most of what a better imputer is worth while spending a third of the error reduction, whereas the same total reduction placed on unimportant features recovers only a fraction. This holds at two missingness rates, on classification, and, with the sign reversed, under informative missingness, where better imputation instead costs tree ensembles. The diagnosis yields a procedure that imputes only where importance sits, at a fraction of the cost when the imputer is expensive and with no loss we can detect. Finally, treating imputer choice as a decision, we find that the diagnostic practitioners can compute is unreliable precisely when the missingness is informative, and is then beaten by a default fixed in advance. Better imputation is not failing to improve reconstruction; the improvement lands where the model cannot use it.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.