Tensor Network Datamodels for Training Data Attribution
Abstract
Datamodels predict how a model’s behavior changes when different parts of its training data are included or removed. However, linear datamodels assign fixed contributions to training units and cannot represent non-additive joint effects. We introduce Tensor Network Datamodels (TNDMs), which use low-rank tensor networks to model dependencies among training units while keeping the representation compact. TNDMs learn a map from subsets of the training data to the outcome of training on those subsets, allowing them to represent interactions among several training units. We prove that complementary and redundant training groups can induce high-degree subset functions with compact low-rank representations. A single fitted TNDM can then be used for several training data questions, including predicting the effect of removing data, measuring the importance of individual training units, identifying interactions between training units, and selecting useful subsets of the training data. We evaluate TNDMs on controlled tasks designed to contain strong interactions, as well as standard fine-tuning settings for classification and reasoning. On a controlled reasoning task, TNDMs outperform every baseline while observing only half as many training subsets, and TNDM-TT reduces prediction NMSE by relative to the linear datamodel after observing approximately of all possible training subsets. In factual-association fine-tuning, a single TNDM-CP raises held-out from to over the linear datamodel. These results demonstrate that compact low-rank models improve prediction when training data effects depend on combinations of training groups.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.