acceptodds
Under review as a conference paper at ICLR 2027

Learning from the Past: Historical Representation Reuse with Negative Feedback for Vision-Language Models

Abstract

Vision-language models typically rely on cross-attention to capture interactions between visual and textual representations, yet intermediate representations generated during training are rarely reused across training units. Moreover, fluctuations in these representations may reduce the stability of cross-modal learning. To address these issues, we propose Historical Intermediate Representation Reuse and a Negative-Feedback Adapter. HIRR reuses the cross-attention Key and Value representations from the previous training unit, allowing historical intermediate information to participate in subsequent cross-modal interactions. The Negative-Feedback Adapter further regulates the current representation according to its deviation from a historical reference, suppressing unnecessary representation fluctuations. We analyze these mechanisms from the perspectives of statistical variance reduction and local perturbation propagation, linking historical representation reuse to effective sample complexity and feedback stability. Experimental results across multiple vision-language models and image-text retrieval benchmarks show that combining historical representation reuse with negative-feedback regulation consistently improves cross-modal learning performance and data efficiency, while demonstrating applicability across different cross-attention-based architectures.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.