Omni-Persona: Systematic Benchmarking and Improving Omnimodal Personalization
Abstract
While multimodal large language models have advanced across text, image, and audio, personalization research has remained primarily vision-language, with unified omnimodal benchmarking that jointly covers text, image, and audio still limited and lacking methodological rigor for absent-persona scenarios and systematic grounding studies. We introduce Omni-Persona, the first comprehensive benchmark for omnimodal personalization. We formalize the task as cross-modal routing over the Persona Modality Graph, encompassing 4 task groups and 18 fine-grained tasks across items. To rigorously diagnose grounding behavior, we propose Calibrated Accuracy (), which jointly evaluates correct grounding and appropriate abstention by incorporating absent-persona queries into a unified evaluation framework. Our experiments reveal three findings: (i) recent open-weight models exhibit a consistent audio-versus-visual grounding gap that RLVR partially narrows through dense rule-based supervision; (ii) recall and parameter scale are incomplete diagnostics, as strong recall can coexist with absent-persona hallucination and larger models do not always achieve higher , establishing calibration as a distinct evaluation axis; and (iii) SFT is constrained by the scalability of ground-truth annotation. Although RLVR improves recall, it does not improve calibrated abstention under our reward design, leaving at or below the base model. Omni-Persona therefore provides a diagnostic framework for identifying pitfalls in omnimodal personalization and guiding future post-training and reward design.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.