Response Is Not Dispensability:A Controlled Study of Counterfactual Data Selection
Abstract
Selecting training evidence from a model's intervention response requires separating response magnitude from learning utility. We examine a low-response retention rule and show that probability saturation favors large-margin identities even at a fixed intervention-induced logit contrast. An exact binary decomposition and a multiclass saturation construction characterize this effect. We test its consequences in synthetic and rendered-digit tasks with known label-preserving interventions, matched class budgets, and shared coverage rules. In a linear factorial, increasing core strength changes rare-evidence retention from to under low-response selection, while a nuisance-logit ranking retains approximately half. Nonlinear experiments on synthetic and two colored-digit tasks separate retained composition from retrained subgroup accuracy. Across all 18 common training settings, applying the intervention during training on the same retained identities improves mean four-group worst accuracy over training without intervention. With validation-selected configurations, the gains are , , and percentage points on synthetic data, colored digits, and colored MNIST. The results establish that response, retention, and intervention use are distinct decisions that require separate validation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.