Personalized but Unchecked: Safe Personalization Requires Conditional Control in Medical Advice
Abstract
Safe personalization requires following user preferences when doing so is safe and overriding them otherwise. In medical advice, however, multiple language models can recognize clinical constraints and follow safe preferences separately, yet fail to combine these abilities reliably. We propose a two-stage explanation: preferences can impair patient-specific judgment formation and prevent correct judgments from constraining treatment recommendations. Repair should therefore support both reliable judgment and constraint-guided action. Controlled dialogues behaviorally localize these failures: preferences introduced before judgment can distort its formation, while those introduced after a correct self-generated judgment can redirect recommendations. Internal analyses and causal interventions link these failures to a shared component carrying clinical information. Switching between clinical-decision and preference-fulfillment framing shifts both judgments and recommendations, and the same clinically relevant component participates in this framing effect at recommendation. Intervening in the component at judgment and recommendation jointly recovers safe advice missed by either intervention alone, supporting complementary repair. Following this direction, training models to follow preferences subject to clinical constraints improves safety while preserving safe preference following, whereas suppressing preferred answers at matched safety incurs greater personalization costs. Benefits extend to single-turn and free-text advice, while further evaluations examine reasoning gains and non-medical capability retention. Together, these findings provide a unified account of safe-personalization failures and a repair principle grounded in conditional control across judgment and recommendation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.