From Preference to Action: A Benchmark for Continual Preference Memory in Multimodal Agents
Abstract
Multimodal agents must learn user preferences from repeated feedback and apply them to future tasks, yet accurate preference recall does not ensure preference-consistent generation. We introduce Pref2Act, a benchmark for continual preference memory that jointly evaluates preference understanding and personalized image generation. We use fine-grained user profiles to construct multimodal histories that interleave dialogue, image generation, and image editing. Throughout these interactions, user feedback progressively reveals, reinforces, and updates preferences. Its evaluation covers explicit fact retrieval, multi-turn information integration, implicit preference inference, while repeated anchor questions measure how agents' understanding of users evolves as interactions accumulate. We evaluate representative multimodal agents under different memory configurations to examine preference retention, adaptation, and application across scenarios. Our results show that current models struggle to retain preferences implicitly revealed through user feedback and do not consistently apply even correctly identified preferences to subsequent image generation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.