VIPSim: A Dual-Role User Simulation and Judgment Framework for Personalized VLM Alignment
Abstract
Personalized vision-language model (VLM) alignment has become an increasingly important as general-purpose VLMs still fall short of satisfying individual needs. However, the user perspective in this area remains predominantly overlooked. Existing approaches confine interactions and evaluation to limited, static, and single-turn settings. They fail to capture the complexity of real-world multi-turn dialogues, where users' needs vary across individuals and continuously evolve throughout conversations. Furthermore, user simulators remain largely unexplored for multi-modal personalized conversational settings. Prior work has shown that simply prompting assistant-aligned models to act as users often produces unsatisfactory and unrealistic behaviors. In this paper, we introduce VIPSim, a novel dual-role simulation and judgment framework designed to bridge the gap between static evaluation and real-world personalized VLM alignment. Conditioned on minimal personalized profiles and visual assets, our simulator dynamically synthesizes realistic, multi-turn interactions centered around personalized entities. It also serves as a judge to provide internal, multi-dimensional, session-level feedback on the assistant's performance in personalized scenarios. Experiments demonstrate that our fine-tuned VIPSim significantly outperforms both open-source baselines and proprietary models in simulating human-like interactions. Notably, our approach also substantially mitigates hallucination and sycophancy compared to its vanilla counterpart. By leveraging personalized user simulation, our method enables evaluation of personalized VLM alignment performance in realistic settings, bringing personalized VLM alignment one step closer to satisfying individual needs.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.