acceptodds
Under review as a conference paper at ICLR 2027

Endpoint Exposure and Constraint Coverage in Mixup under Spurious Correlations

Abstract

Selective Mixup changes both interpolation and endpoint sampling. We study these effects using two clean-endpoint baselines. The policy-matched baseline follows the original pair sampler and coefficient weights, whereas the marginal-matched baseline reproduces pooled group-position exposure with independently sampled endpoint groups. For binary labels and attributes, we define Disentangling Pair Exposure (DPE) as the probability that a sampled pair differs in exactly one variable. DPE and complementary same-polarity coverage jointly determine endpoint-polarity balance. In a linear squared-loss model with lambda drawn from a Beta(alpha, alpha) distribution and finite alpha > 0, Mixup and endpoint training uniquely identify the noiseless core and spurious signal coefficients if and only if both endpoint polarities have positive mass. Positive DPE is sufficient but not necessary for noiseless signal identification, while pair topology can change objective curvature. Experiments on synthetic data, Colored MNIST, biased CIFAR-10, Waterbirds, and CelebA show that endpoint sampling alone can produce substantial changes in worst-group accuracy. With balanced synthetic endpoint groups at source agreement probability rho = 0.98, mean Mixup worst-group accuracy ranges from 90.34% to 90.59% as DPE varies from zero to one. The remaining Mixup–endpoint differences depend on the pair relation and interpolation space. These results motivate evaluating selective Mixup together with clean-endpoint baselines derived from its sampling policy.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.