Know Me To Recommend Me: Personalized Recommendation from Long-Term Human-MLLM Assistant Interactions
Abstract
Imagine that in the near future, when you ask your personal Multimodal Large Language Model (MLLM) assistant, "Recommend a laptop that would suit me," it can recover your preferences from sparse, incidental cues accumulated through everyday interactions and directly identify the right item, without requiring you to repeatedly spell out all details. To bridge this research gap, we propose LIP-Rec, which studies how everyday experience between human and their MLLM assistants can be leveraged for personalized recommendation. To systematically evaluate such capability, we introduce PersonaRec, a large-scale benchmark that lays the foundation for moving personalized recommendation from modeling behavioral signals within recommendation platforms toward a new paradigm grounded in daily interactions with MLLM assistants. It comprises 200 distinct personas and over 30K human-refined candidate items, spanning three common daily-life domains: Fashion, Food, and Movies, and 21 sub-domains. For each persona, we construct 83.5K length daily multimodal interactions, where item preferences are sparsely and naturally revealed through explicit, implicit, and evolving preference patterns, and also be grounded in multimodal clues. We benchmark frontier MLLMs and find that they perform poorly on PersonaRec: they struggle to recover item preferences from long interaction histories and translate them into effective retrieval queries. Therefore, we propose Yo'Rec, a simple baseline method that targets these limitations and enhances open-source MLLMs on PersonaRec. Empirical results show its efficacy, offering a promising direction for LIP-Rec.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.