CoMo: Efficient Video Transformer via Codec Residual and Motion-Field Geometry
Abstract
Transformers have demonstrated remarkable capabilities in video understanding, but their computational cost grows rapidly with the large number of spatiotemporal tokens. Existing token reduction approaches commonly rely on learned selection modules, intermediate model features, or comparisons between spatially co-located patches, introducing additional overhead, requiring redundant tokens to be partially processed, or failing to recognize redundancy under spatial displacement. Video codecs, in contrast, have already characterized temporal redundancy through motion compensation and residual coding. In this paper, we systematically investigate codec-derived residuals and motion vectors as lightweight signals for video token reduction. We find that residuals provide an effective measure of information not explained by motion-compensated prediction, whereas raw motion-vector magnitude captures absolute displacement but overlooks spatial variations in the motion field. Based on these findings, we introduce , a training-free and model-agnostic framework that combines codec residuals with motion-field geometry to adaptively remove motion-compensated redundancy before the vision encoder. Across action recognition and video language understanding benchmarks, CoMo consistently improves the tradeoff between performance and efficiency, reducing VLM computation by up to 53.2% and accelerating inference by up to while retaining accuracy comparable to dense inference.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.