Unlearning Sycophancy Improves Machine Creativity in Large Language Models
Abstract
Large Language Models (LLMs) tend to echo users' preferred answers, leading to the sycophancy issue. This conversational habit inherently stifles divergent thinking and restricts genuine machine creativity. Directly supervising models on creativity-oriented data often causes out-of-distribution degradation, while current sycophancy interventions are effective exogenously or compromise general model capabilities. Instead of teaching models to imitate creativity, we formulate sycophancy mitigation as a behavioral unlearning problem and explore whether unlearning undesirable sycophantic behaviors can provide an alternative pathway for improving LLM creativity. We construct sycophancy-specific forget sets through self-inoculation and introduce a lightweight positional refinement of negative preference optimization. Our method outperforms standard fine-tuning and unlearning baselines in sycophancy mitigation, creativity enhancement, and utility preservation, establishing sycophancy unlearning as an effective, cost-efficient approach to enhance LLMs' creativity without direct supervision.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.