When Properties Are Not Separated Do LLMs Fail at Judging
Abstract
Evaluation of Large Language Models (LLMs) is critical for reliable deployment, and the LLM-as-a-Judge paradigm is widely used for scalable evaluation involving verifiable factual properties and subjective properties. Prior work uses a single judge model or loosely coordinated multi-agent models that produce a holistic score, which empirical works claim to be unreliable under distribution shift and subject to biases such as self-preference and sensitivity to output length; recent works mitigate these issues through prompting, ensembling, or tool augmentation, but retain the same structure by combining verifiable factual properties and subjective properties within a single evaluation process. We hypothesize that evaluation error is governed by verifiable factual properties coverage α, where lower α increases reliance on subjective properties, and propose Judgment Decomposition Graphs (JDGs) to increase by representing evaluation as a directed acyclic graph over evaluative properties in four phases: Phase I extracts and normalizes properties, assigns verification types , and constructs ; Phase II applies executable verifiable factual properties and subjective properties according to ; Phase III aggregates node-level outputs into a scalar score; and Phase IV calibrates the score as a function of verifiable factual properties coverage . Expected evaluation error is bounded by , where denotes executable verification reliability. Across ReviewMT and JudgeBench, JDGs achieves HA of 0.82/0.88, reduces SPB by up to 46.6%, and lowers CE by up to 43.1% over baselines. Increasing verification coverage improves HA from 0.69 → 0.83 and 0.75 → 0.89, while iterative optimization reduces the proxy gap from 0.20 → 0.03 and robustness drift from 0.124 → 0.061.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.