Two-Phase Caching and Activation for Input-Efficient Fine-Tuning of Vision Transformers
Abstract
Following the success of the ViT, transformers have been widely adopted for image recognition, multimodal understanding, and visual generation, increasing the need for efficient fine-tuning. Although parameter-efficient fine-tuning (PEFT) reduces trainable parameters, it can still require back-propagation through the full token sequence, retaining substantial activation and computation costs. Token pruning and merging address this overhead by shortening the sequence, but also modify the context available during forward propagation. Building on selective token back-propagation, we formulate input-efficient fine-tuning (IEFT) for vision transformers to reduce token-level backward computation while retaining full-sequence forward context. To enable token selection informed by the completed forward pass, we propose a two-phase caching-activation (TCA) strategy that first caches global keys and values during a gradient-free full forward pass, then recomputes selected tokens against this cached context. With fixed parameters and consistent deterministic computation, this replay reproduces the corresponding full-sequence forward outputs while restricting gradient propagation to selected token paths. Experiments across 13 recognition tasks demonstrate over 60% peak-memory reduction and approximately 45% training-time reduction with minimal accuracy degradation on ViT fine-tuning. Combining TCA with PEFT yields additional efficiency gains. Extensions to Qwen3-VL-8B, FLUX.2-4B, and Wan2.2-14B cover multimodal understanding and visual generation, yielding 3.4× average micro-batch capacity for the first two and approximately 2×video generation training throughput on the same hardware. The code is available in the supplementary materials and will be publicly released.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.