acceptodds
Under review as a conference paper at ICLR 2027

AltRA-Test: An Alternative Annotator Test for LLM Judges Without Shared Annotators

Abstract

Large language models (LLMs) are increasingly used as judges and graders to replace human annotators, raising the question of when such replacement is statistically justified. To support such decisions, the Alternative Annotator Test (alt-test) was introduced for dense settings where all items are annotated by the same annotators. However, in many large-scale annotation settings, items are typically annotated by different subsets of annotators. We therefore propose the Alternative Annotator Test with Resampled Annotations (AltRa-Test), a statistical test that does not require shared annotators across all items. AltRa-Test uses a bootstrap that resamples items, draws human annotations for each item, and substitutes one of them with an LLM annotation. This produces paired all-human and mixed human–LLM panels. A replacement is supported when the inter-annotator agreement (IAA) of the mixed panel is statistically non-inferior to the IAA of the all-human panel. We first establish that AltRa-Test is a consistent test of asymptotic level . We then evaluate it through finite-sample size simulations and real-world experiments in which we test 18 LLMs and 5 prompting techniques on 20 datasets spanning four modalities: text, image, audio, and video. Across these evaluations, AltRa-Test requires as few as 100 items to yield stable replacement decisions. In simulations, it dominates baselines by improving the trade-off: for any fixed rate of falsely accepting models noisier than simulated humans, , it correctly accepts more models that are less noisy. In real-world data, AltRa-Test supports replacing humans with LLMs in about half of the evaluated setups.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.