acceptodds
Under review as a conference paper at ICLR 2027

MergeNet: Spatial Routing and Execution Trade-offs in Learned Token Compression

Abstract

Vision transformers preserve fine image detail but pay to process every patch through global blocks. Learned token merging must decide where information flows while producing a shorter sequence for later computation. We introduce MergeNet, a local-to-latent transformer that combines differentiable spatial mass transport with an exact-budget gather in one forward pass. Routing redistributes features and mass on the original patch grid; selected carriers then read full-grid features through cross-attention before latent encoding. Mass-aware latent attention retains the transported token mass, and an exact query/key augmentation implements its bias in an unmodified FlashAttention call. On ImageNet-1K with DeiT-S/8, MergeNet reaches 81.360% top-1 with 392 final patches and 82.618% after 512-pixel fine- tuning. Its 224-pixel configuration halves the latent patch sequence and reduces estimated multiply–accumulate operations by 37.0% relative to dense DeiT-S/8. In a separate synthetic A100 benchmark, compiled inference takes 27.24 ms at batch 64, compared with 27.22 ms for compiled ToMe. Evaluations across token budgets and image resolutions characterize the accuracy and implementation cost of this trainable physical bottleneck.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.