Greedy MMD Without Replacement: Selecting and Stopping for Mixture Separation
Abstract
This paper studies how to split an unlabelled sample drawn from a mixture when a separate reference sample from is available. Greedy maximum mean discrepancy (MMD) selection without replacement chooses distinct observations to approximate , leaving a remainder intended to approximate . Success depends on order and stopping: early choices may remove observations needed later, while a wrong stopping size can leave a poor remainder even when the selected observations fit . An error bound tracks the effect of earlier selections. Under two geometric conditions and appropriate stopping, convergence rates are established for the MMD errors of the selected group relative to and the remainder relative to . Simulations and real-data experiments illustrate the effect of stopping and support the predicted convergence behaviour.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.