acceptodds
Under review as a conference paper at ICLR 2027

Linearizing Vision-Language ViTs Beyond the Training Resolution

Abstract

Linear recurrent visual encoders offer increasing computational advantages at high resolutions, yet representative linear CLIP models exhibit substantial performance degradation when the visual sequence length moves beyond their low training resolution. Strong pretrained vision-language ViTs retain high-quality representations beyond their training resolutions, whereas naively linearized recurrent encoders fail to inherit this capability. In this work, we introduce LADDER (**L**ayer **AD**aptive **D**istillation into **E**fficient Hybrid Encode**R**), a lightweight framework that converts pretrained vision-language ViTs into hybrid recurrent encoders while preserving their cross-resolution capabilities through architectural designs. LADDER combines attention and recurrence within blocks and across layers. For global feature aggregation, our proposed Vision-Delta pairs recurrent patch mixing with Attention-based CLS (AttnCLS), decoupling global readout from the block's recurrent propagation. For dense feature interaction, Resolution-Robust Layer-Adaptive Hybridization optimizes full-attention placement under a fixed budget via reinforcement learning guided by cross-resolution teacher patch-feature recovery. We primarily evaluate LADDER on TIPSv2, with additional validation on SigLIP2. With only 22 epochs of image-only training on ImageNet-1K, LADDER closely matches or surpasses its teachers across vision-language and linear probing evaluations, while matching the ViT teacher to within 1% point at Large scale. On high-resolution Cityscapes inputs, LADDER achieves the throughput of the FlashAttention teacher. Code and models will be released.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.