Towards Personalized Visual Imagination from Everyday User–MLLM Interaction with Unified Multimodal Models
Abstract
Text-to-image (T2I) models has achieved remarkable progress in visual quality and even physical-world reasoning. Nevertheless, when facing personalized image-generation requests that specify only what users currently care about most, the unspecified visual spaces in their resulting images remains shaped by population-level priors, imagining an "average user" rather than the particular individual behind the prompt. Meanwhile, as MLLMs serve as personal assistants, their everyday experiences with users naturally disclose how each individual tends to imagine these open visual spaces. Therefore, we propose VIMG, a new image-generation task that studies how these user–MLLM histories can foster personalized visual imagination for users' daily painting requests. Accordingly, we propose VIMG-Bench to evaluate VIMG, spanning six daily image-generation domains and 30 sub-domains, with a three-level hierarchy of visual imagination preferences (global, domain and sub-domain). It features 100 diverse users, each paired with 168K multimodal interactions with their MLLM assistant that may occasionally reveal such preferences, as well as 3K personalized T2I requests. We benchmark mainstream T2I models and Unified Multimodal Models (UMMs), and find that although UMMs excel at reasoning over complex generation instructions, they remain limited on VIMG, which poses the opposite challenge: imagining what is left unspecified from simple yet highly personalized image requests. We further propose VIMGer, which reinforces UMMs for this new imagination-to-realization paradigm, providing a foundation for advancing the VIMG.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.