ZipSign: Training-Free Spatio-Temporal Token Compression for Sign Language Translation
Abstract
Sign language translation (SLT) models have made steady progress by improving visual representation learning and translation modeling. However, existing SLT models usually process redundant visual tokens across spatial and temporal dimensions, leading to substantial computational costs and limiting their deployment in resource-constrained and real-time scenarios. In this work, we propose ZipSign, to the best of our knowledge, the first training-free spatio-temporal token compression method for sign language translation. ZipSign adopts a two-stage token compression strategy. The first stage performs similarity-based spatial token merging, which preserves sign-critical hand and face regions while merging redundant background tokens within each frame. The second stage performs confidence-guided temporal token pruning, which removes less informative frame-level visual tokens before they are fed into the translation module. ZipSign introduces no additional learnable parameters and can be directly plugged into existing SLT models during inference. Extensive experiments on standard SLT benchmarks demonstrate that ZipSign effectively reduces computational cost and inference latency while maintaining translation performance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.