acceptodds
Under review as a conference paper at ICLR 2027

EUDAIMONIA: Evaluating Undesirable Dynamics in AI

Abstract

Large language models (LLMs) are increasingly used as conversational partners for companionship, emotional disclosure, and interpersonal advice, but the social dynamics of these interactions can create harms that are not captured by current evaluations in real-world settings. We introduce the Social AI Design Code, a set of principles for keeping LLMs from encouraging harmful intimacy, dependence, or prolonged engagement, which we developed with input from psychologists and trust-and-safety practitioners. To evaluate these risks in natural and diverse user–LLM interactions, we operationalize the code with , a benchmark of 969 opening-turn queries and 3,147 design-requirement violation checks built from WildChat through weak-to-strong filtration, multi-model relabeling, and controlled rewriting. Evaluating 26 recent LLMs, we find rapid but uneven progress: within 19 months, Grok shifted from the highest to the lowest violation rate, with Grok-4.7 violating only 13.8% of checks. Although Claude-Opus-5.5 and GPT-6-Astra outperform Grok-4.7 on general benchmarks, they violate more checks, showing that greater capability does not guarantee adherence to our design code. Extended thinking does not reduce violation rates, suggesting these failures are social-alignment problems rather than deficits solvable through test-time reasoning alone. \reframe{Finally, we evaluate models on 229 multi-turn WildChat continuations, violation rates increase, but cross-model trends mirror those in opening turns, showing that our findings generalize to multi-turn interactions.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.