acceptodds
Under review as a conference paper at ICLR 2027

Safety is Contextual, LLM-Judges are Not: Navigating the Rigid Priors of Evaluators

Abstract

LLMs-as-judges are the primary way to evaluate safety at scale. Despite their importance, LLM-judges themselves are rarely evaluated beyond human agreement in simple, static benchmarks. We therefore investigate two under-explored but crucial properties of LLMs-as-judges: their **sensitivity** to in context-information, and their **steerability** to differing safety definitions, which may not align with their internal safety priors. We evaluate the safety judging abilities of 13 generalist LLMs and safety-specific judges, and investigate the impact of novel in-context information and changing safety definitions. We find that while LLM-judges can learn from new information, they are broadly unlikely to adjust their evaluations if the context or safety definition contradicts their prior.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.