COERCION:Benchmarking Moral Coercion in Dialogues via Psychological Dimension Decoupling
Abstract
Moral coercion is a form of implicit manipulation that uses shared ethical norms to exert interpersonal pressure. While existing NLP models detect explicit toxicity, they struggle with moral coercion due to its polite phrasing and contextual dependency, Current datasets on moral coercion are insufficient, often relying on movie scripts that lack the pragmatic complexity of actual social interactions. To address this gap, we build COERCION, a dataset of 4,700 multi-turn dialogues sourced from social media. To model the latent manipulative mechanism of moral coercion, we design a taxonomy decomposing moral coercion into four quantifiable psychological dimensions: Obligation, Constraint, Value Judgement, and Toxicity. We further propose a two-stage dimension-aware pipeline that extracts these dimensions as structural anchors, guiding downstream Large Language Models (LLMs) toward deeper pragmatic reasoning rather than shortcut learning. Statistical analysis indicates that our dataset reflects moral coercion characteristics. Experimental results demonstrate that grounding LLMs in these dimensions improves detection accuracy. Furthermore, ablation results indicate that the joint integration of ValueJudgement and Constraint is sufficient to sustain near-optimal performance across models, demonstrating their critical contribution to the exploration of moral coercion.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.