PersonaTrust: Benchmarking Trustworthy Personalized MLLM Assistants
Abstract
As MLLMs become lifelong personalized companions, the reliability of their customized responses assumes unprecedented importance: users expect them not only to provide highly tailored assistance, but also to honestly acknowledge when user-specific evidence and information are insufficient, thereby avoiding misleading or unsafe recommendations. To bridge this gap, we propose PersonaTrust, the first benchmark dedicated to trustworthiness throughout MLLM personalization. In contrast to existing MLLM trustworthiness benchmarks that mainly assess model reliability on questions with well-defined answers, PersonaTrust focuses on personalized questions with inherently open-ended answer spaces. Therefore, trustworthiness is evaluated by whether the model can faithfully narrow this space according to the available personalized evidence. We observe that existing frontier MLLMs and personalization methods fall short on PersonaTrust, particularly by creating a false sense of familiarity with users, e.g., providing generic recommendations without clarifying that they are not grounded in sufficient personalized evidence. To remedy this, we propose a simple plug-and-play method, PR-Calibrator, which equips these assistants with two-layer calibration safeguards before and after the original pipeline, improving reliability while largely preserving their personalization capabilities. Experiments on PersonaTrust demonstrate that PR-Calibrator clearly improves the trustworthiness of existing personalization methods, while also highlighting the need for future MLLM assistants to better recognize the boundaries of their user-specific knowledge.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.