acceptodds
Under review as a conference paper at ICLR 2027

When Is Gradient-Based Data Selection a Property of the Predictor?

Abstract

Gradient-based data selection is commonly interpreted as identifying examples that are informative for a given predictor, yet neural-network parameters are not unique representations of the predictor. We study predictor-level well-definedness of the resulting discrete selection under function-preserving reparameterizations. Across BADGE, CRAIG, and GradMatch, equivalent predictors replace – of a query batch while their logits remain unchanged to numerical precision; in end-to-end LoRA fine-tuning, factor-coordinate selection replaces – of MiniLM queries and, in the analytical factor-basis comparison, on average for ModernBERT. Per-example gradients transform as covectors, so a fixed parameter-space metric generally generally changes the geometry seen by the selector. We characterize when Gram geometry is preserved exactly or up to scale, and give selector-specific constructions in which the resulting distortion crosses a discrete selection boundary. For Fisher-based selection, we further show that a coordinate-fixed prior can break invariance even when the Fisher terms transform covariantly. Finally, repeated acquisition reveals that query non-identifiability and downstream utility are distinct: acquisition trajectories can diverge strongly while terminal accuracy remains similar in the settings studied. These results identify parameterization as a source of data-selection non-identifiability that is separate from conventional questions of subset utility.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.