ConceptTune: Tuning Concept Vectors for Generalizable Video Understanding
Abstract
Despite their impressive capabilities, multimodal large language models (MLLMs) still face challenges in video understanding, which current approaches mainly address through parameter fine-tuning, represented by LoRA and LOFIT, or activation steering, represented by RepE. The former enables accurate task-specific adjustment by updating model weights or parameter offsets, but may alter pretrained capabilities. The latter offers flexible generalization through direct representation-space intervention, but their positive–negative activation contrasts may not accurately represent the target concept. We propose ConceptTune, a hybrid method combining the accuracy of fine-tuning with the generalization ability of activation steering. ConceptTune fine-tunes a learnable concept offset directly in model activations rather than model parameters, treating this offset as the concept. Each concept offset is learned from many examples through concept-targeted generative completion. This training procedure encourages the concept offset to capture semantic information shared across training examples of the same concept type, yielding a reusable concept representation that can be applied as a lightweight, plug-and-play activation intervention at inference time. Experiments show that ConceptTune improves transfer across video-grounding benchmarks and compares favorably with parameter-efficient fine-tuning and contrastive activation steering, supporting concept-targeted learning in activation space.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.