acceptodds
Under review as a conference paper at ICLR 2027

COERCION:Benchmarking Moral Coercion in Dialogues via Psychological Dimension Decoupling

Abstract

Moral coercion is a form of implicit manipulation that uses shared ethical norms to exert interpersonal pressure. While existing NLP models detect explicit toxicity, they struggle with moral coercion due to its polite phrasing and contextual dependency, Current datasets on moral coercion are insufficient, often relying on movie scripts that lack the pragmatic complexity of actual social interactions. To address this gap, we build COERCION, a dataset of 4,700 multi-turn dialogues sourced from social media. To model the latent manipulative mechanism of moral coercion, we design a taxonomy decomposing moral coercion into four quantifiable psychological dimensions: Obligation, Constraint, Value Judgement, and Toxicity. We further propose a two-stage dimension-aware pipeline that extracts these dimensions as structural anchors, guiding downstream Large Language Models (LLMs) toward deeper pragmatic reasoning rather than shortcut learning. Statistical analysis indicates that our dataset reflects moral coercion characteristics. Experimental results demonstrate that grounding LLMs in these dimensions improves detection accuracy. Furthermore, ablation results indicate that the joint integration of ValueJudgement and Constraint is sufficient to sustain near-optimal performance across models, demonstrating their critical contribution to the exploration of moral coercion.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.