acceptodds
Under review as a conference paper at ICLR 2027

SeqGuard: Length-Independent Stopping for Streaming Hallucination Detection

Abstract

Streaming detectors can score a language model's output for hallucinations while it is being written. The rule for when to interrupt is still a tuned score threshold, and a threshold has no error control that holds at every length: set to on RAGTruth, it reaches on long outputs, and calibrated on AggreFact or TofuEval and deployed on RAGTruth it runs over budget. We introduce SeqGuard , a stopping rule that turns any streaming detector's scores into an anytime-valid sequential test. Each score is ranked against scores from known-faithful outputs, the rank becomes a bet, and the output is interrupted once the accumulated winnings reach . A martingale argument bounds the probability of ever interrupting a faithful output by , for any detector and any length. The proof covers a conditionally calibrated variant that keeps of the deployed rule's power; the simpler marginal variant we deploy is checked empirically. With a prefix-entailment scorer and a hidden-state probe, the operating configuration's realized false-alarm rate stays at or within its confidence interval of the budget in-domain (its one point estimate above budget is Data2txt at , of episodes) and under shift among RAGTruth, AggreFact, and TofuEval, helped by slack (in-domain FAR is of ). A far-out-of-distribution HaluEval test breaks both rules: SeqGuard overshoots on summarization and on dialogue, against and for the tuned threshold. Against a memoryless rule that sees the same two detectors, the power gain is concentrated on Data2txt (– on three of four annotation-defined classes at matched FAR) and at on QA and Data2txt; it does not win on Summary. The cost is delay ( tokens at ), though – of flags fire before the output ends.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.