Mitigating Conformity in Large Language Models via Social-Context-Aware Unlearning
Abstract
Humans often revise their judgments when surrounded by a confident majority. Large Language Models (LLMs), increasingly used as conversational assistants and collaborative agents, exhibit a similar tendency and may answer correctly in isolation but switch to an incorrect answer after observing peer responses or discussion history. Prior works have explored mitigation through prompt-level interventions and reflection. However, these approaches mainly reshape the input context at inference time. This raises a natural question of whether conformity can be mitigated through parameter-level correction rather than prompt-level intervention alone. Our key observation is that conformity-inducing interaction traces naturally reveal both what should be forgotten and what should be preserved. Based on this observation, we propose to mitigate conformity through machine unlearning. Intuitively, unlearning updates model parameters to reduce the likelihood of socially shifted responses under conformity-inducing contexts, while retaining supervision helps preserve reference-supported behavior. We develop a conformity-targeted unlearning framework that identifies answer shifts, constructs forget-retain pairs, and applies selective unlearning to suppress conformity-driven errors without erasing task competence. Our experimental results show that unlearning can reduce socially induced errors while preserving the model's original reasoning ability. These findings suggest a new perspective on LLM conformity: rather than being solely a test-time prompting issue, conformity can be viewed as a behavioral failure mode that can be explicitly identified and corrected through selective unlearning. Our code is available at https://anonymous.4open.science/r/xdefewgferh.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.