BiCLoRA: Two-Sided Constrained LoRA for Audio-Video-Text Continual Learning
Abstract
Pretrained audio-video-text models align complementary auditory, visual, and linguistic information within a shared representation space, enabling unified recognition of real-world events. In deployment, these models must continually acquire concepts beyond their pretraining data, often without retaining past observations because of privacy and storage constraints. We formulate this problem as exemplar-free Audio-Video-Text Continual Learning (AVT-CL), which extends an aligned multimodal space while preserving prior knowledge and cross-modal alignment. We identify two failure modes of continual low-rank adaptation: *historical interference*, where the input factor overlaps with activation subspaces occupied by previous sessions, and *alignment drift*, where a freely trained output factor writes new information into arbitrary shared-space directions, even when the input side is protected. To address both, we propose *Bilaterally Constrained Low-Rank Adaptation* (**BiCLoRA**), which fixes both factors before optimization. A history-safe input basis derived from a bounded eigensummary of past projection inputs restricts where updates may act, while a frozen output basis spanning the leading directions of current-session representations restricts where new information may be written. Only a small coupling kernel for each modality is optimized and subsequently merged into the current projection. Across three pretrained AVT backbones and three benchmarks, BiCLoRA improves macro final average accuracy from to and backward transfer from to over the strongest matched adaptation baseline. Ablations confirm that the input constraint protects historical representations, the output constraint preserves cross-modal alignment, and effective retention requires both.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.