Beyond Pointwise Attribution: Relationship-Aware Training Data Attribution
Abstract
Training data attribution is commonly formulated by assigning each training example a scalar influence score for a target model behavior. This formulation treats influence as an intrinsic property of an individual sample, yet model parameters are learned by jointly optimizing over the entire training distribution. As a result, the contribution of one example is inherently conditional on the presence of others: perturbing one sample can change not only its own effect, but also how other samples contribute to the learned model. Existing pointwise attribution methods collapse these coupled effects into marginal scores and cannot explicitly represent substitution, reinforcement, or dependency among training examples. Motivated by this, we reformulate training data attribution from estimating isolated sample importance to characterizing how training examples influence one another’s contribution to a target behavior. Instead of asking only which samples are important, we ask how the importance of one part of the training data changes when another part is perturbed. This leads to a novel ***Relationship-Aware Training Data Attribution*** task that captures compensatory, reinforcing, and suppressive dependencies among training samples. Specifically, we organize training samples into coherent groups according to their attribution-relevant representations and construct a query-conditioned asymmetric attribution graph over these groups. Each node represents a coherent region of the training data, while each directed edge measures how perturbing one group changes the attribution assigned to another. Since inter-sample effects are inherently asymmetric, we explicitly preserve their directionality rather than treating relationships as reciprocal. By modeling training influence at the level of structured and asymmetric group interactions rather than isolated examples, our approach provides a more faithful characterization of how different parts of the training data collectively shape model behavior.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.