Knowledge Networks Enable Stable and Effective Attribution: Data Shapley Valuation in One Training Run
Abstract
In-Run Data Shapley makes principled data valuation feasible for modern foundation models by estimating sample contributions along one observed training trajectory. However, existing in-run estimators mainly treat examples as independent contributors, making attribution scores unstable under noisy gradient signals and prone to favoring semantically redundant samples. A direct second-order treatment could model these interactions, but Hessian-dependent computation is costly and difficult to scale. To address these challenges, we propose Graph-Aware In-Run Data Shapley, a simple but effective approach that introduces a sample-level knowledge network into one-run valuation. The network connects semantically related examples and is used to smooth structurally supported contributions while penalizing redundant neighbors during selection. This design stabilizes attribution, improves diversity-aware data valuation, and replaces expensive Hessian-based interaction modeling with sparse neighborhood aggregation. Extensive experiments on image and language learning tasks show that the proposed graph-aware estimator improves attribution stability, subset selection, and computational efficiency over state-of-the-art approaches.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.