acceptodds
Under review as a conference paper at ICLR 2027

MUSE: A Multilevel User Feedback Learning Framework for Self-Enhancement in Personalized Multimodal Agents

Abstract

Multimodal agents must produce reliable outputs across modalities and adapt to individual users. A single overall score, however, gives little guidance on whether an error stems from one modality, a conflict between modalities, or a mismatch with user preferences. We introduce MUSE, a Multilevel User Feedback Learning Framework for Self-Enhancing Personalized Multimodal Agents. Within an interaction, Fine-Grained Unimodal Verification (FUV) identifies local defects and constraint violations. Cross-Modal Collaborative Verification (CCV) checks those findings against the combined evidence and identifies conflicts to guide revision. Across interactions, Feedback-Driven Personalization Verification (FPV) learns how verified output properties relate to each user's satisfaction. It updates user-specific adapters from feedback while the task models remain frozen. We evaluate each method on 2,709 instructions completed by 34 participants, yielding 92,106 participant–instruction encounters across 18 task categories and text, visual, and audio inputs. Experiments show that hierarchical verification and feedback-driven adaptation improve multimodal reliability, correction effectiveness, and preference-aligned personalization.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.