When Personalization Compromises Truth: Joint Control of User Alignment and Factuality in Large Language Models
Abstract
Personalization makes large language models more useful by adapting responses to user specific context, but it can also degrade factual reliability when personal signals override world knowledge. We study this failure mode, termed personalization-induced hallucination (PIH), as a dual-objective control problem involving personalization and factuality. We posit that personalized information contains both helpful components that improve user alignment and harmful components that interfere with factual reasoning. To address this challenge, we propose PerFact, an inference-time steering framework that jointly controls the two objectives within their shared activation subspace. PerFact identifies representations relevant to both personalization and factuality, removes PIH-related components from the personalization direction, and combines the resulting safe personalization direction with a factuality-enhancing direction for model intervention. Experiments across multiple backbones and benchmarks show that PerFact achieves consistent gains in both personalization and factuality, mitigating the trade-off between the two objectives observed in existing baselines. Our results identify a characteristic representation shift associated with PIH when incorporating personalized context, and show that targeted representation steering can reduce such shifts while improving factual reliability.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.