acceptodds
Under review as a conference paper at ICLR 2027

DarkPattern Teaming: Adaptive Discovery of LLM Dark Patterns in Multi-Turn Conversations

Abstract

LLMs can use manipulative or deceptive tactics, termed dark patterns, even while helping users pursue legitimate goals. Such dark patterns often surface only after several turns and depend on the user’s needs and circumstances, so single-turn benchmarks and adversarial red teaming miss many of them. We introduce DarkPattern Teaming, a framework that simulates multi-turn conversations across 24 combinations of user needs and everyday contexts, detects dark patterns turn by turn, and uses the outcomes to select which conditions to simulate next. On Qwen 3.5 27B, 44.9% of detected conversations show their first dark pattern only after the opening response. Under equal simulation budgets, adaptive sampling yields 22.1% more flagged conversations than balanced stratified sampling, and blinded annotators confirm 90% of sampled detections (compared with 72–86% for baselines). Across five open- and closed-weight LLMs, DarkPattern Teaming reveals differences in the types of manipulation detected among models with similar discovery yields. The discovered failures also provide mitigation data: we pair manipulative responses with revisions that remove manipulation while preserving help with the user’s goal. Activation steering with 100 pairs reduces Qwen 3.5 27B’s held-out dark-pattern rate from 42.2% to 13.0%.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.