Image Robustness Does Not Imply Video Robustness: The Limits of Single-Frame Evaluation
Abstract
Single-image corruption benchmarks such as ImageNet-C are a cheap way to choose a robust backbone for a video system, but this shortcut works only if image and video corruptions rank the models the same way. We test how well a model's robustness to single-image corruptions actually predicts its robustness to video corruptions. We evaluate models on image and video corruptions under identical conditions and check whether the single-image robustness ranking predicts the video robustness ranking, using a rank-correlation threshold for reliable prediction. Our central finding is that image robustness transfers to video only under certain corruption conditions. Specifically, we observed that the single-image ranking transfers to video to the extent that a single frame shows the corruption. When a corruption realization is shared across a video clip, the single-image ranking is preserved almost perfectly across seven settings. By contrast, when a corruption disrupts the temporal structure of a clip, the single-image ranking does not reliably predict the video ranking in any setting. These temporal-structure corruptions include frame drop, repetition, and reordering, which damage the frame sequence even though every delivered frame stays clean. Corruptions that a single frame shows only in part give mixed results, for example when the corruption strength drifts over time or the camera shakes. Single-image robustness had essentially no rank correlation with video robustness to temporal-structure corruptions among 37 video-pretrained 3D backbones. The same pattern was observed in codec and transport pipelines: Single-image robustness predicts robustness to codec compression well, but poorly predicts robustness to packet loss or variable-frame-rate judder. Our theoretical analysis proves that two corruptions can look identical to a single-image benchmark yet reverse the ranking of two video models, implying that image-only testing cannot guarantee the right model choice for video. We recommend adding a separate targeted temporal evaluation whenever temporal structure can be disrupted, and selecting video models by single-image robustness only when a single frame shows the expected corruptions.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.