acceptodds
Under review as a conference paper at ICLR 2027

HALT: Held-Out-Source Evaluation for Hallucination Detection in Forensic Video Reporting

Abstract

Hallucination detectors for long-form multimodal generation are almost always evaluated on text from the same source LLMs they were trained on. In forensic video reporting we show this conflates detection with identifying which LLM wrote the text. A detector that silently stops working when the underlying model is updated offers false assurance, in a setting where inventing a detail and omitting a crime are two different evidentiary failures. We build HALT (Hallucination Across LLMs and Techniques) on a previously released corpus of 19,361 forensic video reports (21.8M words) in which three frontier multimodal LLMs each describe all 807 surveillance videos, spanning eleven crime categories, under all eight prompting strategies, labelled for six hallucination types by a cross-judging panel in which no model evaluates its own output. Our central result is that much of what a detector appears to know is the identity of the model that wrote the report: a predictor reading no report text at all, returning only the writing model’s base rate, captures 70 to 73% of a detector’s above-chance signal on the fabrication labels against 13% on omission. Withholding a source LLM accordingly costs the fabrication labels 0.104 to 0.191 AUC against 0.035 to 0.105 for omission and distortion, consistently across four detector families from bag-of-words to a fine tuned long-context encoder, and again when a source LLM is updated rather than replaced, though the tendency being detected is unchanged across that update. Two further findings concern the panel’s H1–H6 labels: (i) report length predicts the panel’s scene-fabrication label at AUC 0.727 but the consensus of two human annotators who each independently labelled the same 130 reports at only 0.566, and (ii) models are not allowed to judge their own output, so each source LLM is scored by a different pair of judges; those pairs differ in how often they agree, so comparing hallucination rates across models partly compares their judges. We release the reliability tiers, the four splits, and reference implementations of every detector.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.