BiMetaJudge: Bidirectional Metamorphic Auditing and Selective Risk Control for LLM Judges
Abstract
Large language models (LLMs) are increasingly used as judges. Repeated sampling estimates local decoding variability, and position swaps expose one presentation artifact, but neither establishes whether a judge tracks a known change in answer quality. We present BiMetaJudge, a black-box audit built around a graded bidirectional response curve: matched nuisance controls estimate spurious movement, while independently verified mild/severe degradations and partial/full repairs measure directional elasticity. A development-fitted risk model drives a staged policy whose score contains only purchased probes, and a disjoint split certifies the final accepted set. Stable-but-wrong cases are an evaluation slice rather than a claimed new phenomenon. We evaluate the complete pipeline on MetaJudge-MR, spanning four domains, five judge configurations, and 12,000 base comparisons. On the held-out test split, BiMetaJudge reaches AUROC 0.824 versus 0.808 for the strongest non-BiMeta baseline and stable-but-wrong recall 0.795 versus 0.280. Leave-one-domain and leave-one-judge gains remain +2.67 and +2.68 AUROC points over nuisance-only evidence.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.