acceptodds
Under review as a conference paper at ICLR 2027

PROVENANCE NETWORKS: END-TO-END EXEMPLARBASED EXPLAINABILITY

Abstract

We introduce provenance networks: a shared backbone feeds a task branch (e.g. classification) and a jointly trained index branch that predicts each input’s exemplar index within the training set, making the network behave like a learned, end-to-end k-nearest-neighbor classifier. Training-set index accuracy alone is a poor measure of usefulness: a model reaching 96–98% index accuracy on MNIST can still fail almost completely at recovering a known source for a perturbed test input. We introduce a test-time, ground-truth-backed source-recovery protocol (MNIST, FashionMNIST, CIFAR-10) that exposes this gap and show it closes once the index branch is trained with augmentation matched to the evaluation perturbations – at which point it decisively beats an embedding-space kNN baseline, most dramatically on CIFAR-10 (70.8% vs. 2.2% top-1 under the strongest perturbation). A ProtoPNet-style baseline is structurally incapable of this task, capped near 0% by construction despite comparable classification accuracy. We further stress-test a proposed use case, dataset debugging via index-branch entropy, against synthetic label noise, and find the opposite result: entropy performs at chance for mislabeled-data detection. On scalability, a subset of training exemplars preserves task accuracy with little loss, and a factorized index head cuts the index layer’s parameters by 41–500× but converges substantially more slowly. Together, these results argue for evaluating exemplar-attribution methods against test-time, ground-truth-backed protocols rather than training-set accuracy alone.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.