Navigating User Behavior toward Personalized Multimodal Generation
Abstract
Modern AIGC pipelines deliver high-fidelity images and videos but presuppose a well-formed creation instruction, while end users rarely articulate visual details, so personalized content generation must turn a user's interaction history into an executable instruction that aligns downstream synthesis with user demand. Serializing behavior is largely settled—discrete item identifiers from generative recommendation already provide a serviceable interface—so the open obstacles lie in learning: the desired instruction is unobserved and underdetermined, and no signal ties a written instruction back to user intent. We therefore propose NaviGen. An evolution-guided induction phase manufactures its own supervision, using evolutionary search to discover history-to-instruction traces from which the model distills preference reasoning and instruction writing. A closed-loop alignment phase then optimizes a triangular self-consistency reward that couples the written instruction, the model's own preference prediction, and the target semantics, so instruction quality earns credit only when it is user-grounded. Experiments across diverse domains show that NaviGen improves personalized image and video generation and yields more specific, relevant, and visually generatable instructions, with next-item prediction reported as an intermediate diagnostic. Our code is released at: https://anonymous.4open.science/r/NaviGen-08F8/.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.