Certified Data Attribution After Model Reselection
Abstract
Data deletion can change the configuration selected by validation as well as its fitted parameters. We study attribution to the resulting retrain-and-reselect procedure over a finite menu of strongly convex learners. A selected output can be certified without identifying the winning configuration. We characterize the feasible attribution set under interval information with deterministic ties and derive its minimax absolute-error radius. Residual regions and validation-conditioned prediction bounds give an adaptive procedure whose refinement bound depends on disagreement among near-optimal candidates. A three-observation ridge construction has a nonzero deletion effect and an output-certification cost independent of an arbitrarily small validation gap under a shared damped-Newton schedule. Full Newton steps remove this separation, showing that refinement granularity matters. Across 4,800 classification queries, reselection occurs in 12.3–19.1% of queries by setting. Output stopping saves no refinements at a joint signed-logit tolerance. A complete probability-output sweep yields 1.5–12.4% fewer refinements than equally certified winner identification for individual targets, with substantially smaller gains for simultaneous outputs. Additional conditioning contributes little and does not improve measured latency. The results characterize when the output, tolerance, and solver permit computational savings.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.