GTR: Gated Token Recurrence for Efficient Dense Prediction
Abstract
Self-attention-based vision backbones perform well on dense prediction, but the quadratic computational cost of global softmax attention limits their efficiency as image resolution increases. We introduce **Gated Token Recurrence (GTR), a softmax-free recurrent vision backbone that combines gated linear attention**, alternating spatial scan directions, and spatially enhanced SwiGLU blocks. GTR is distilled from a detection-specialized DINOv3 teacher using only final-layer patch-token alignment through a linear projection and squared loss, without masked-token prediction or intermediate-layer supervision. With Objects365 detector pre-training, GTR-L achieves 58.9 box AP on COCO val2017 with 1.908 ms median batch-one latency under compiled FP16 execution on an RTX4090. The same backbone also transfers to instance segmentation, pose estimation, oriented detection, semantic segmentation, and monocular depth estimation. For efficient inference, we develop a specialized chunkwise CUDA operator that is faster than FLA v0.5.0 at 1.6K tokens on RTX4090. TensorRT deployment on DRIVE AGX Thor achieves 2.282–8.769 ms median batch-one latency across the evaluated models. These results show that recurrent token mixing can provide an efficient alternative to global softmax attention for high-resolution dense prediction and edge deployment. Code and pretrained models will be released upon publication.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.