acceptodds
Under review as a conference paper at ICLR 2027

AR-Retrofit: Retrofitting Pretrained Decoders with Attention Residuals for Adaptive Depth

Abstract

Attention Residuals improve Transformer performance by enabling adaptive aggregation of representations across different depths. However, applying this mechanism to existing pretrained models remains challenging, as direct integration requires expensive backbone retraining and introduces additional inference overhead. To avoid costly backbone retraining, we propose Attention Residual Retrofit (AR-Retrofit), a lightweight framework for frozen pretrained decoders. AR-Retrofit introduces depth aggregation through gated residual corrections while preserving the original forward function at initialization. The magnitude of these corrections is adaptively controlled by Residual Strength Modulation for different tokens and depths. To reduce the additional inference overhead, we propose Routing-guided Residual Skipping (ReSkip), which reuses the routing signals learned by AR-Retrofit for selective layer execution without training an additional controller. Experiments on language, multimodal, and robotic tasks show that AR-Retrofit transfers adaptive depth aggregation to pretrained models while improving performance across all three domains. On Qwen3-VL-2B, it adds less than 0.4% trainable parameters and requires only 12.6 minutes of training on a single H100. Across six VLM benchmarks, average scores improve by 2.08 and 1.44 points for Qwen3-VL-2B and 4B, respectively. ReSkip increases decoding throughput over Base by 4.5%–8.8% while retaining most gains. These results show that our framework enables lightweight and efficient adaptive depth in pretrained models.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.