How One Training Example Changes Others: Exact Route Attribution for BatchNorm
Abstract
In models with batch normalization, an example affects learning through both its own loss and the shared statistics used by other examples. Training-data attribution measures how individual examples influence a model, but a total influence score does not explain how that influence arises. We introduce Edge-Game Replay (EGR) to separate these contributions as they propagate through training. EGR intervenes on loss participation, normalization statistics, later batch membership, and running-buffer updates, and uses exact Shapley values over complete training replays to allocate the final evaluation-loss difference. We find that shared statistics and loss participation make contributions of comparable magnitude. Across 16 CIFAR-10 ResNet runs, adding statistic contributions to the loss contribution reduces mean absolute prediction error by 66.9% after 393 updates and raises Kendall correlation from 0.391 to 0.822 on a disjoint evaluation set. The resulting score ranks deletions more accurately than TracInCP, last-layer influence functions, and single-model TRAK at all three comparison horizons. Eight CIFAR-100 WideResNet runs also show improved two-update prediction. The decomposition also guides the choice among loss masking, exclusion from normalization statistics, and example deletion.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.