Training Witnesses: Trusting the Training without Trusting the Trainer
Abstract
Progress in machine learning cannot outpace our ability to verify it. With an explosion in papers today, every scientific claim rests initially on trust in the trainer, leading to uneven evaluation, baselines, and forestalling of reliable progress. Traditionally, the burden of verification falls on the reader, who must reproduce expensive training runs. This strategy is impractical due to an explosion in slop contributions, diversity of methods, and the sheer compute required. We put the burden of proof where it belongs, on the trainer, and in the process also cut the overall cost of verification significantly. We introduce **Witnesses**, a method for certifying training data usage and evaluation in a neural network training run. We leverage the insight that fast behavioral fingerprints with occasional replay challenges is sufficient for auditing neural network training. Our method is applicable at scale with low overhead to the trainer, is cheap for the verifier, rejects bad training runs with amplifiable probability, and allows for exact queries of both data inclusion and exclusion. We test our method through baseline language model training runs ranging over 100M to 2B scales, across DDP and FSDP, and demonstrate minimal overhead. We also introduce a self-regulating leaderboard of training runs that enables shared baselines and progress. We invite the community to participate in the leaderboard to improve reproducibility in machine learning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.