Mask-space distributionally robust optimization for learning under stochastic missingness
Abstract
Training with stochastic feature masking enables classifiers to handle missing data during inference. However, standard masked empirical risk minimization (ERM), even when aligned with the evaluation missing-rate distribution, can still overfit, leading to degraded performance across varying missing rates. We propose mask-space distributionally robust optimization (MaDRO), a transport-cost DRO framework centered at the joint data-mask distribution induced by a nominal masking rule, with transport restricted to mask space and toward additional missingness. MaDRO has a closed-form robust objective function: the base masked loss plus a hinge penalty on the worst deletion-cost-adjusted loss increase caused by one additional feature deletion. For each masked training example, MaDRO either keeps the sampled mask unchanged or deletes one additional observed feature, choosing the deletion only when its loss increase outweighs its deletion cost. Across public tabular and vision benchmarks, controlled semi-synthetic studies, structured missingness shifts, and attribution-fidelity analyses, MaDRO improves nominal-mask performance and reduces overfitting relative to aligned masked ERM.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.