Moral Mazes in the Era of Large Language Models
Abstract
Navigating complex social situations, from giving critical feedback without hurting morale to rejecting requests without alienating teammates, is an important part of corporate life. Although large language models (LLMs) are entering the workplace, it is unclear how they will navigate or even reshape these norms. To investigate this question, we created HR Simulator, a game where users play as an HR officer and write emails to tackle challenging workplace scenarios, evaluated with GPT-4o as a judge based on scenario-specific rubrics. We analyze over 600 human and LLM emails and find systematic differences in writing style: LLM emails are more formal and empathetic, and human emails are more varied. Across four LLM judges, human emails receive lower average pass rates than LLM-only emails (8-44% versus 27-68%). Rewriting human drafts with an LLM helps mainly when that LLM is a weak writer on its own: rewrites by Claude 3.5 Haiku and Gemini 2.5 Flash beat those models' own emails in five of eight judge comparisons, compared with three of twelve for stronger writers. On the evaluation side, across 23 judge models, judges show more agreement with one another on emails' outcomes as their capability increases. A deeper analysis of 10 judge models' email rankings shows that specificity and empathy are broadly rewarded, and more capable judges weight them less. Scenario-specific tact has a smaller effect that appears to grow with judge capability. Together, our results suggest that as LLMs increasingly become readers of professional emails, human writing may fail to meet their standards, and human-LLM co-writing may better match their preferences.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.