BOTTLENECKED SEMANTIC COMPRESSION BY REPRESENTATION RECONSTRUCTION
Abstract
Vision Transformers incur high computational costs when processing long visual token sequences. Existing reduction methods typically rank tokens independently and discard those deemed unimportant or redundant. We propose a Reconstruction-guided framework for visual token condensation(ReConV), which that summarizes the complete visual sequence with a compact Summary Bank (SMB) and retains tokens whose information is not adequately captured by the SMB, as indicated by their reconstruction residuals. At each compression layer, a lightweight position-conditioned decoder reconstructs the current tokens, leaving shared information compressed in the SMB while explicitly preserving poorly reconstructed details. Experiments on image classification and multimodal understanding demonstrate favorable performance–efficiency trade-offs: ReConV surpasses the uncompressed baseline with \(48%\) fewer GMACs on ImageNet and enables reusable image-level representations for efficient multi-turn inference. Code is available at TokenDrop.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.