Learning from the Past: Historical Representation Reuse with Negative Feedback for Vision-Language Models
Abstract
Vision-language models typically rely on cross-attention to capture interactions between visual and textual representations, yet intermediate representations generated during training are rarely reused across training units. Moreover, fluctuations in these representations may reduce the stability of cross-modal learning. To address these issues, we propose Historical Intermediate Representation Reuse and a Negative-Feedback Adapter. HIRR reuses the cross-attention Key and Value representations from the previous training unit, allowing historical intermediate information to participate in subsequent cross-modal interactions. The Negative-Feedback Adapter further regulates the current representation according to its deviation from a historical reference, suppressing unnecessary representation fluctuations. We analyze these mechanisms from the perspectives of statistical variance reduction and local perturbation propagation, linking historical representation reuse to effective sample complexity and feedback stability. Experimental results across multiple vision-language models and image-text retrieval benchmarks show that combining historical representation reuse with negative-feedback regulation consistently improves cross-modal learning performance and data efficiency, while demonstrating applicability across different cross-attention-based architectures.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.