SMILE: Small-Large Collaboration for Efficient Long-Term MLLM Personalization
Abstract
Enabling general-purpose MLLMs to adapt to each user's unique traits, habits, and evolving preferences is essential for building next-generation personalized multimodal assistants. Existing methods for personalizing MLLMs into lifelong human companions often adopt an ahead-of-time, query-agnostic context management strategy. They pre-compress and condense interaction histories into memories after each interaction, risking the loss of personalized multimodal evidence required in subsequent user queries. To overcome this limitation, we posit that an MLLM should browse the complete raw interaction history just-in-time conditioned on each incoming personalized query, and introduce SMILE, a novel SMall-large collaboration framework for effIcient Long-tErm MLLM personalization. SMILE breaks "one-model-fits-all" paradigm, specializing a Small MLLM (S-MLLM) as a lightweight yet intelligent personalized context distiller and employing a Large MLLM (L-MLLM) as a powerful but expensive personalized context reasoner. Specifically, given a contextualized and idiosyncratic user query, the S-MLLM learns to distill compact personalized context from the long raw interaction history, and hand it off to the frozen L-MLLM for reasoning and response generation, delivering efficient and effective customized assistance through collaboration. Experiments on four diverse benchmarks show that SMILE achieves superior performance while handing off averagely about 2.6% of the interaction history to the L-MLLM. Moreover, since SMILE keeps L-MLLM untouched during training, it enables nearly "free-lunch" scaling of personalized reasoning at inference time by seamlessly plugging in larger and stronger L-MLLMs.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.