VLoRA: Adapting Vision-Language Models for Multimodal Recommendation via Two-Stage Low-Rank Adaptation
Abstract
Learning item representations that unify content semantics and collaborative filtering (CF) signals is crucial for sequential recommendation. While recent large language model-based methods have shown promise, they often ignore visual information that is important for user decisions. Vision-language models (VLMs) provide a natural foundation for multimodal recommendation, but directly fine-tuning them poses three challenges: objective interference between CF alignment and embedding learning, modality-length asymmetry, and representation collapse under sparse supervision. We propose VLoRA, a parameter-efficient two-stage post-training framework that adapts a frozen decoder-only VLM into a CF-aware multimodal embedding model. In the first stage, collaborative signals are injected through autoregressive next-item prediction over user interaction sequences. In the second stage, item embeddings are refined using bidirectional attention and three complementary objectives: Multimodal Masked Next Token Prediction for modality-asymmetric masking, dual-view self-contrastive learning for discriminative embeddings, and representation-preservation regularization to prevent representation collapse. With progressive LoRA unfreezing, VLoRA introduces trainable parameters equivalent to only 0.4% of the backbone. Experiments show that VLoRA consistently outperforms twelve strong baselines in both in-domain and out-of-domain settings, achieving an average improvement of 4.4% on NDCG@10 across all evaluation settings. Further analyses verify the effectiveness and generalizability of the proposed framework.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.