acceptodds
Under review as a conference paper at ICLR 2027

Ruleusion: Generator-Free Metamorphic Testing and Answer-Conditioned Probes for Hallucination Detection

Abstract

Without a reference answer or access to a model's internals, a hallucination has to be detected from the model's own replies. Sampling detectors ask the same question many times and read disagreement as a warning. When every sample is the same, the agreement can mean the model is right or that it repeats one error; we call this decoding collapse, and these detectors then have little to measure. Detectors that change the input instead, such as SAC³ and MetaQA, pay a large language model (LLM) to write the tests. We present Ruleusion, which checks an answer by asking questions whose answers follow from it by logic. Each test is either a fixed rewrite of the question or a short follow-up about the model's own answer. The rewrite states a logical relation, such as negation, presupposition, or entailment, so the expected answer is known before the model is queried, and the reply is graded locally. No generative model writes or grades any test. On answers that five LLMs gave to questions from five short-answer datasets, Ruleusion achieves the highest average AUROC of seven detectors, 0.744, ahead of the strongest baseline, SAC³ (0.722), and it is the most accurate detector at every call budget: with 6.5 calls per question it already matches SAC³ at 27 calls. It uses 54% fewer calls and 70% fewer tokens than SAC³. On the half of answers where all twelve samples coincide, every other detector loses accuracy, SelfCheckGPT and Semantic Entropy falling to 0.53–0.56 AUROC, while Ruleusion keeps 0.69, the same as on the other half.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.