WorldWeaver: Benchmarking Multimodal Intent Recognition and User Modeling in Interactive Visual Generation
Abstract
Personalized interactive continuous image generation is an important task for adaptive visual information distribution, in which visual content evolves with user intent and preferences. Existing benchmarks typically study final-image personalization, multimodal intent recognition, and user modeling separately; they rarely represent generated visuals as environments for later feedback or preserve alternative continuations from a shared history. We introduce WorldWeaver-Bench, an extensible benchmark for this setting. Its shared-prefix branching trees support joint evaluation of multimodal interaction-intent recognition, persistent user-profile modeling, and matched comparisons of alternative paths. Experiments show that story context improves the interpretation of current interaction intent, longer and broader interaction evidence improves profile recovery, and preference-compatible paths are not consistently the most diagnostic for user modeling. WorldWeaver-Bench provides a testbed for studying how generated visual content both responds to user preferences and reveals them through interaction.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.