acceptodds
Under review as a conference paper at ICLR 2027

Rethinking the Evaluation of Training Data Attribution in Generative Models

Abstract

Concerns about the misuse of copyrighted and private data in training image generation models motivate training data attribution (TDA), which assigns each training sample a score measuring its influence on a generated query. TDA methods are commonly evaluated using the Linear Datamodeling Score (LDS), which aggregates attribution scores over sampled subsets to predict changes in model behavior after retraining. However, attribution scores exhibit a long-tailed magnitude distribution, with most scores close to zero. Finite sampling introduces noise into their estimation, and repeated estimates show substantially greater sampling variation relative to the observed score energy in low-influence regions. These fluctuations reflect noise-induced errors in measuring weak influences. Consequently, changing aggregation weights alone can alter LDS without changing the underlying attribution scores, undermining the reliability of method comparisons. For example, restricting the contribution of low-influence regions improves prediction of the full subset responses. Power transformations such as squaring can also improve prediction by suppressing the contribution of low-magnitude scores, giving methods an evaluation advantage without improving their attribution estimates. We propose Signal-to-Noise Ratio LDS (SNR-LDS), which models measurement noise as zero-mean Gaussian and estimates a signal-to-noise ratio for each attribution score. By excluding scores below a common SNR threshold, SNR-LDS limits the contribution of unreliable measurements and aims to reduce the evaluation advantage conferred by reweighting. Deletion experiments on CIFAR-2 Flow Matching models show that removing influencers identified by the method ranked highest by SNR-LDS increases mean query loss more than removing those identified by the method ranked highest by LDS. We further introduce controlled-source retrieval, which directly evaluates source rankings by recovering injected training samples from foreign concepts, complementing SNR-LDS. Code and Data are available at https://anonymous.4open.science/r/SNR-LDS-65B2.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.