The Evaluator Is a Random Variable: Auditing the Fixed Feature Space Behind Text-to-Motion Benchmarks
Abstract
Text-to-motion generation is benchmarked almost entirely through one artifact: a text-motion contrastive feature extractor released in 2022 whose embedding space defines FID, R-precision, matching score, Diversity, and MultiModality on HumanML3D and KIT-ML. Every confidence interval in the literature is computed over generation repeats with this evaluator held fixed, so it omits the evaluator's own training randomness. We treat the evaluator as a random variable. Retraining it from its public recipe with 30 independent seeds on KIT-ML, we score frozen pools of generations from four released state-of-the-art models under every seed and decompose metric variance into evaluator-seed, generation-seed, and batching components. The evaluator seed is the largest component for FID, R-precision, and matching score: for MoMask, the FID standard deviation across seeds is 0.05, seven times the batching noise that published intervals report and larger than the generation-seed component. Among the orderings of our pools under the released checkpoint, every pair whose seed-mean gap is within two seed standard deviations reverses under some seeds, and the gap under the released checkpoint does not predict which; a non-generative baseline that returns the nearest training motion for each test caption overtakes the strongest generator under half of the seeds, an advantage that disappears when duplicate retrieved motions are removed. Inference hyperparameters selected on one evaluator transfer to other seeds with regret no larger than the seed noise, while the identity of the optimum does not. A corruption suite shows that FID ignores a foot drift larger than the standard deviation of its own velocity features, and identity controls show that the feature-extraction route alone can move FID by more than a published state-of-the-art margin. We release the seeded evaluators, the frozen pools, and an evaluator-adjusted confidence interval so that future improvements can be reported against the noise floor of the measurement itself.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.