Penalized-Weight Robust Optimal Transport: From Hard Clipping to Soft Compression
Abstract
Robust optimal transport (OT) methods reduce the influence of outliers, and several produce point-level outlier flags, but we know of no guarantee that these flags recover the contaminated observations. We introduce PW-ROT (Penalized-Weight Robust OT), which assigns each transport path a weight and penalizes linearly. Eliminating the weights in closed form yields ordinary OT with a bounded ground-cost family , indexed by a threshold and a power : is ROBOT's hard-clipped cost (up to its conventional factor of two in the threshold), and gives smooth soft-compression costs. In a general parametric model, we show that the bounded cost down-weights separated outliers automatically: although no data are discarded and the outlier fraction is not an input, the PW-ROT criterion matches, up to a constant, a criterion that removes that fraction of mass from the fitted model (not from the data), exactly for the hard clip () and up to an explicit error for . We bound the estimation error under arbitrary outliers, prove consistency of the hard clip with a fixed fraction of separated outliers in bounded-support models, and show that the flags recover the outlier set exactly with high probability whenever the estimation error is small relative to a separation margin. For observations with Gaussian inliers in of variance and a point target, exact recovery holds for every member, including ROBOT, with a fixed fraction of outliers separated at the same scale, when ; slower growth provably flags inliers for any estimator. The power governs a trade-off: soft powers down-weight rather than discard outliers beyond the threshold, so their weights vary continuously with the cost, but separated outliers then bias the estimate toward themselves, which the hard clip avoids. Empirically, hard clipping is best for one-sided, well-separated outliers; soft compression has lower bias near the threshold and more stable threshold selection on images; and is the most consistent default.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.