acceptodds
Under review as a conference paper at ICLR 2027

RAIM: Robust Aggregation of Inexpensive Models for Hallucination Detection

Abstract

Automatic evaluation of faithfulness increasingly relies on a large language model acting as a judge, yet the most reliable judges are proprietary frontier models. Their per-call cost, rate limits, and off-premises inference, however, make them ill-suited to the continuous, high-throughput monitoring that benchmarking and production use demand. In this work we investigate three fundamental questions: whether a panel of cheap open-weight judges (4–9B parameters) can be aggregated to stand in for a frontier one, what the substitution sacrifices in agreement (Cohen's ) and accuracy, and when it is worth making. To address them, we propose RAIM, an aggregation scheme robust to the members' correlated errors, which couples: () a cross-fitted stacked logistic regression that fits the members' weights jointly, therefore penalising weaker judges echoing a stronger one; and () an admissibility test that, read from the members' own outputs, identifies a regime in which aggregating them improves on their best member and stays within reach of the frontier judge. We instantiate RAIM under ten judges drawn from disjoint families and we measure the substitution across eight faithfulness benchmarks. Read as paired differences against Claude Sonnet, the substitution clearly improves on one benchmark and clearly worsens on three (only two by a non-negligible margin), leaving the remaining four unresolved. The panel retains a median 93% of the frontier judge's agreement and gives up only 2.9 points of balanced accuracy on average. Inference runs at a sixty-fourth of the frontier's price at conventional cloud rates, or at the electricity cost on premises, so the operative expense is a one-time in-domain calibration of some fifty to a hundred labelled records, amortised over a few tens of thousands of evaluated items. Against the literature, the panel is competitive with purpose-trained detectors on their home benchmarks, losing 1.3 points of accuracy to GPT-4o and 1.9 to the best LLM-AggreFact leaderboard model; yet the strongest detector we reran falls 6 points behind on our grounded sets, and cannot score the ungrounded ones. In conclusion, whether pursuing the aggregation is worthwhile depends on the members themselves, making the gain auditable from how widely competence is spread across them and how far their errors decorrelate — both read at no further cost off the aggregator's calibration set. If several capable members err on different items, the panel improves on its best judge and approaches the frontier; conversely, where one dominates, the stacker recovers the leader, yet only there does the frontier remain materially ahead. A panel of cheap judges can therefore stand in for a frontier one at a small fraction of the cost wherever this audit admits it.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.