acceptodds
Under review as a conference paper at ICLR 2027

RhetoriFact: Decomposing Belief Robustness and Comparator Fragility in Grounded LLM Judges

Abstract

Recently, Large Language Models (LLMs) have demonstrated remarkable capabilities as automatic evaluators for fact-checking. However, existing LLM judges remain susceptible to stylistic biases, often favoring authoritative rhetoric over factual substance. This raises a fundamental question: whether these biases are intrinsic to LLM evaluators or artifacts of ungrounded evaluation designs. To systematically investigate this vulnerability, we formalize factuality judgment as a decomposition of a belief function (discriminating truth) and a comparator function (recognizing factual equivalence), and introduce RhetoriFact. Our benchmark comprises 2,445 validated claim quadruples systematically permuted along a 2×2 factorial of factuality and rhetorical presentation. Extensive evaluations across eight state-of-the-art LLMs reveal the dual nature of rhetorical contamination: while the belief function is robust—models overwhelmingly resist direct falsehoods when anchored with oracle evidence—the comparator function fails pervasively, frequently flipping verdicts on factually identical pairs due to pure rhetorical variation. Logit-level analysis and scoped-prompt ablations further confirm that this instability stems from model probability distributions rather than sampling noise. Finally, we demonstrate that incorporating Chain-of-Thought (CoT) reasoning can partially mitigate these failures, though no output-surface predictor fully secures the rankings. Our findings conclude that while evidence grounding secures individual factuality verdicts, it does not eliminate the stylistic biases that skew automated rankings.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.