Do Frame-Order Differences Between Video Pretraining Stages Survive Fine-Tuning?
Abstract
Linear probes on frozen features are a standard way to compare what different pretraining gives a video model. We ask whether the differences they reveal survive fine-tuning, for frame order, measured as the balanced-accuracy drop when a clip's frames are shuffled. Across the published pretraining stages of one architecture (VideoMAE), on two binary fall-detection datasets, one synthetic and one in the wild, frozen checkpoints differ markedly at the output, and supervised Kinetics training lowers the fall probe's order sensitivity even as it raises accuracy on the synthetic dataset. Fine-tuning collapses most of these differences. On the synthetic dataset the fraction of the gap between the self-supervised (MAE-only) and Kinetics checkpoints that fine-tuning closes is κ = 1.08 [0.84, 1.33], so the gap closes completely within error. It also closes under a standard fine-tuning schedule (κ = 0.85 [0.22, 1.36]), under re-drawn and tubelet-preserving shuffles and with nonlinear probes, and it does not survive on native-resolution input or with cross-fitted probes, where refitted probes even find the fine-tuned order reversed. On the in-the-wild dataset, where κ is uninformative, the spread between checkpoints shrinks by 69%. Between these two checkpoints, what closes is how the fall read-out uses frame order, not the linearly readable direction information the output carries: on a task defined by frame order their frozen outputs are not distinguishable, and both kept most of their direction information through fine-tuning for fall detection. In exploratory layer-wise probes, the frozen checkpoints differ mainly in the upper layers, and under the standard schedule they still differ at blocks 9 and 10 yet agree at the output. In TimeSformer, whose frozen Kinetics output carries no linearly readable direction information, the fall-task gap does not close under either recipe (κ = 0.08 and 0.43, both intervals containing 0), and most of the gap in direction information persists; the same holds in Hiera, where we predicted before fine-tuning that its fall-task gap would not close (κ = 0.03 [−0.54, 0.45]). On an order-defined task TimeSformer's gap closes (κ = 0.86 [0.69, 1.01]) and Hiera's, as the probe reads it, does not (κ = 0.25 [−0.24, 0.60]). Within VideoMAE, fall-task gaps closed even where differences in direction information remained. A frozen comparison of frame-order sensitivity on a task that does not require order thus describes, for VideoMAE, mostly how a probe reads the top of the network, and it cannot tell which differences fine-tuning will keep.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.