acceptodds
Under review as a conference paper at ICLR 2027

SAVOR: Sample-Aware Video Token Reduction with Native Variable-Length Execution

Abstract

Video Transformers spend much of their computation on temporally redundant tokens: nearby frames are often very similar, and static backgrounds do not change across frames. Many token-reduction methods either keep the same number of tokens for every sample, or keep masked tokens in the computational graph during training. The former ignores per-sample redundancy, while the latter does not shorten the actual sequences that are fed to attention layers. To address these issues, we propose Sample-Aware Video tOken Reduction (SAVOR), which operates on a variable-length transformer backbone. It flattens all clips in a batch into one sequence with cumulative lengths and uses a per-sample attention mask to prevent cross-sample attention, so that the removed tokens can be physically deleted during both training and inference. SAVOR combines three components: a deterministic PreMerge that merges redundant temporally adjacent tokens, a CLS-conditioned scorer that keeps the highest-scoring tokens, merges low-scoring neighbors, and drops the rest, and a Hierarchical FLOPs Allocator that turns a global FLOPs target into per-sample, per-layer retention ratios. We evaluate SAVOR on two video benchmarks, i.e., EPIC-KITCHENS-100 (EK-100) and Kinetics-400 (K400). On EK-100 multi-instance retrieval, SAVOR reduces the analytical encoder FLOPs proxy by 30.3%, achieves 60.74 mAP versus 60.71 for the dense backbone, and improves nDCG by 0.37 percentage points. Across the EK-100 target sweep, the realized proxy ratio differs from the target by less than 2% in relative terms. Code is released at https://anonymous.4open.science/r/SAVOR.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.