acceptodds
Under review as a conference paper at ICLR 2027

Your Agent Passed the Ethics Exam. It Won't Use It on the Job.

Abstract

Large language models increasingly do professional work as agents, and the ethical content of that work is rarely signposted. Most existing evaluations ask about ethics directly, so they measure what models can do when asked rather than what they do when the ethics is left unflagged. We introduce a benchmark for unprompted ethical engagement in realistic agentic work, built from 150 practitioner-authored GDPval tasks, each run under seven conditions that vary how its ethical content presents. Nine models attempt every scenario as sandboxed agents, and three LLM judges score each run on a seven-dimension rubric validated against 1,444 human raters. Unprompted ethical engagement is rarely visible: on the baseline tasks the strongest of the nine models scores 38.5 on the 0–100 scale, six sit at or below 20, and half of runs proceed on missing information without acknowledging it. Pooled across the nine, scores rise from 19.2 at baseline to 52.5 when the same tasks are enriched so that their ethical content is explicit, while additions designed to be ethically inert move them by 0.5 points and strictly content-preserving cue removal produces no detectable change. Seven frontier models evaluated baseline-only span 23.5 to 58.2, three above the nine-model roster's average under the explicit conditions. The pattern is consistent with models failing to recognise and surface ethical content rather than lacking the reasoning, though the design does not isolate that reading. The corpus, per-judge scores and transcripts will be released on acceptance.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.