LiveHumanRightsBench: Stress-Testing Reliability of LLM Judgments on Live Human-Rights Cases
Abstract
As LLMs increasingly inform high-stake decisions that affect people's fundamental rights and welfare, it becomes critical to examine the reliability of their judgments when case information is presented differently or when they are challenged by adversarial opinions. We introduce LiveHumanRightsBench, a continuously refreshing benchmark grounded in long European Court of Human Rights (ECtHR) case records. For each benchmark instance, we embed an automated perturbation pipeline to test robustness of model judgments when the case information is paraphrased or summarized. We further stress-test if models can hold their stance when facing various adversarial opinions. Our results show that while frontier LLMs have become remarkably more robust than earlier models under summarization and paraphrasing, they still remain highly sensitive to very simple challenges involving naive dissent and social pressure, even when those challenges introduce no new evidence. Claude-Opus-4.6 and GPT-5.6-Sol reverse over 90% of their initial model judgments within two turns of simple adversarial opinion challenges. Strikingly, adding a challenger identity cue sharply lowers these reversals for a lawyer, and even more so for an AI safety researcher. So far we've only observed this intriguing role-cue effect in Opus-4.6 and 5.6-Sol and it remains absent on other highly capable reasoning models such as DeepSeek-V4-Pro. In conclusion, our findings revealed substantial volatility in frontier LLMs that could risk instability and critically impact human rights at scale, calling for greater effort towards more trustworthy AI for decision support.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.