Max-pooling Network Revisited: Analyzing the Role of Semantic Probability in Multiple Instance Learning for Hallucination Detection
Abstract
Hallucination detection is essential for reliable large language models (LLMs). Hybrid methods such as HaMI combine semantic consistency with internal states through multiple instance learning (MIL), but repeated sampling and semantic comparisons impose substantial computational overhead. We analyze HaMI's decision margins and identify conditions under which semantic-probability scaling enlarges the average logit margin. Motivated by this analysis, we revisit classical sentence classification from a local margin dynamics perspective: a shared MLP extracts token-level features, feature-wise max pooling aggregates them, and a linear head scores the response. The detector requires no additional sampling or semantic consistency computation. Across five benchmarks and three LLMs, it is competitive with semantically weighted HaMI in detection AUC while sustaining several thousand QA detections per second. Its larger gains over fixed-position probes on LongFact support aggregating evidence across token positions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.