acceptodds
Under review as a conference paper at ICLR 2027

Training Witnesses: Trusting the Training without Trusting the Trainer

Abstract

Progress in machine learning cannot outpace our ability to verify it. With an explosion in papers today, every scientific claim rests initially on trust in the trainer, leading to uneven evaluation, baselines, and forestalling of reliable progress. Traditionally, the burden of verification falls on the reader, who must reproduce expensive training runs. This strategy is impractical due to an explosion in slop contributions, diversity of methods, and the sheer compute required. We put the burden of proof where it belongs, on the trainer, and in the process also cut the overall cost of verification significantly. We introduce **Witnesses**, a method for certifying training data usage and evaluation in a neural network training run. We leverage the insight that fast behavioral fingerprints with occasional replay challenges is sufficient for auditing neural network training. Our method is applicable at scale with low overhead to the trainer, is cheap for the verifier, rejects bad training runs with amplifiable probability, and allows for exact queries of both data inclusion and exclusion. We test our method through baseline language model training runs ranging over 100M to 2B scales, across DDP and FSDP, and demonstrate minimal overhead. We also introduce a self-regulating leaderboard of training runs that enables shared baselines and progress. We invite the community to participate in the leaderboard to improve reproducibility in machine learning.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.