SleekAware: Benchmarking Social, Ethical, Empathetic, and Cultural Awareness of LLMs
Abstract
Large Language Models (LLMs) are increasingly used in education and healthcare, where they influence human experiences and decisions, and as judges of other models' responses. These roles require adherence to Social, Legal, Ethical, Empathetic, and Cultural (SLEEC) norms. Yet it remains unclear whether LLMs can reliably recognize and evaluate compliance with these norms. We introduce SleekAware, a benchmark for evaluating LLMs' SLEEC reasoning in natural multi-turn conversations. It comprises three parts: (i) real-world Reddit conversations across therapy, PhD supervision, and elderly care, (ii) expert-defined SLEEC rules for each domain, and (iii) over 6,800 crowdsourced human annotations assessing conversation compliance with these rules. SleekAware evaluates LLMs along two dimensions: (i) how closely their final judgments match human judgments, and (ii) how closely the conversation excerpts they select to support their decisions match those selected by humans. Together, these dimensions go beyond agreement with human judgments to probe the evidence underlying LLM decisions, providing a deeper assessment of SLEEC-awareness. We evaluate five state-of-the-art LLMs using strategies ranging from zero-shot to structured reasoning prompts. Models show only partial SLEEC-awareness, struggling with implicit normative reasoning and emotionally nuanced interactions. While some identify rule compliance well, none reliably ground their judgments in conversational evidence. To advance LLMs' SLEEC capabilities, we release our benchmark and annotation framework at https://sleech.net/benchmark.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.