acceptodds
Under review as a conference paper at ICLR 2027

Mitigating Deceptive Behaviour in Large Language Models through Alignment and Unlearning Interventions

Abstract

As large language and agentic models become increasingly intelligent, their deceptive behaviours have begun to pose crucial alignment challenges that underlie extensive societal and economic harms, particularly under conflicting objectives and in high-stakes tasks. A fundamental question is hence whether and how such behaviours can be mitigated or unlearned, and how this can be achieved effectively. In this paper, we provide a comprehensive study through a systematic investigation of a wide range of alignment and unlearning interventions across diverse models, evaluating honesty, over-refusal, and general capability. Our results show that deceptive behaviours cannot be fully characterized by a binary honesty outcome. Model responses under conflicting objectives were categorized into direct refusal, direct compliance, and qualified compliance, with substantial variation across models and interventions. Motivated by this heterogeneity, we introduce Response-Adaptive Alignment (RAA), an alignment method that adapts its optimization objective to the model's observed response behaviour. We demonstrate that RAA achieves the highest mean aggregate honesty on four of the five models while largely preserving general capability. On these four models, RAA also yields lower over-refusal than supervised fine-tuning and preference optimization. Comparatively, we find that most conventional alignment interventions offer robust mitigation, whereas the effectiveness of unlearning techniques is highly model-dependent. Finally, pooled category rates show that most standard interventions reduce direct compliance more strongly than qualified compliance, whereas RAA reduces both at similar rates.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.