Your "Clean Split" May Not Be Clean: Measuring Recording-Level Leakage in Video Segmentation Benchmarks
Abstract
Video-derived segmentation benchmarks are routinely split at frame or clip level, so near-copies of one recording fall on both sides of the split. That this leaks information is known; we ask whether the standard way of measuring it is itself reliable. The conventional comparison subtracts scores from different test sets, confounding the memorisation gain with test-set difficulty. We propose two tools. Protocol-matched evaluation re-scores both protocols' checkpoints on a common test set, separating the test-set swap from the contrast between the two training protocols without retraining: on infrared, visible and a second public benchmark the matched gain is +6.4 to +8.8 IoU, of which the conventional comparison captures 36%, 68% and −17%. Near-duplicate auditing against adjacent-frame and random-pair references tests whether a "clean" split is clean: that benchmark's recording-disjoint split is as duplicated as adjacent frames (72.5% against 70.6%), which is consistent with why the conventional measurement fails there, while the primary benchmark's split passes on both bands (0.5% visible, 6.9% infrared). On the common clean test set the visible-band night score rises from 0.73 to 0.91, and comparisons on the protocols' respective test sets reverse the observed modality ordering. We release the audit and protocol-matched evaluation code with every per-architecture result.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.