acceptodds
Under review as a conference paper at ICLR 2027

From Personalization to Paternalism: Measuring, Locating, and Mitigating Paternalism in LLMs

Abstract

Large language models (LLMs) deployed as personalized assistants are conceptual analogues to recommender systems, raising the question of whether they inherit two well-known profile-driven failures, namely collapse of output diversity and substitution of profile-based content for query-relevant content. We formalize this class of behaviors as LLM Paternalism: the unconsented imposition of the model's inferred judgment about the user on the user's available choices. Paternalism takes two forms. Restrictive paternalism silently removes options the user did not ask to have removed; active paternalism injects profile-driven content the user did not request. We introduce two corresponding deterministic metrics. DivScore measures restrictive paternalism as the relative shrinkage of effective option diversity (via Vendi Score) when a user profile is added. DriftScore measures active paternalism as the fraction of response items that are entailed by the profile but not by the query (via NLI). Both metrics operate on free-form text without a fixed catalog or LLM-as-judge, and reach substantially higher agreement with human annotators and roughly 2× greater reproducibility than LLM-as-judge baselines. We evaluate on our benchmark, a paired with-/without-profile dataset of 2,000 real-user scenarios from Reddit across ten high-stakes domains (e.g., healthcare, finance). Across eight frontier LLMs, paternalism is universal: seven of eight show positive DivScore, all eight show positive DriftScore. To understand where in the model these failures arise, we apply layer-wise linear probing and causal tracing. The two modes localize to distinct internal mechanisms. DivScore corresponds to hidden-state geometry and DriftScore to attention routing, mapping one-to-one onto the two failure modes. Building on these probes, a detection-and-steering intervention reduces DivScore by 30–100% and DriftScore by 35–68% within a 5% perplexity budget.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.