HAMLET: Hallucination Assessment in LLMs for Multi-party, Long-horizon, Evolving meeting Threads
Abstract
Existing meeting benchmarks primarily evaluate models on a single meeting, treating each meeting as a self-contained document. This misses a key property of real-world meetings: facts and decisions evolve across a series of meetings for one project, where a proposal may be introduced, revised, and eventually adopted or remain unresolved. Hallucinations in this setting can mislead users and carry incorrect information into subsequent meetings, yet existing benchmarks provide limited analysis of such hallucination behaviour. We address these gaps with HAMLET (Hallucination Assessment in LLMs for Multi-party, Long-horizon, Evolving meeting Threads), which treats evolving discussion threads as the unit of evaluation. HAMLET extracts threads from AMI and ICSI meeting series and generates deterministic probes from their evolving states. Candidate answers are human-verified against the full series transcript, and a two-stage judge evaluates both reference-level correctness and unsupported content against the transcript. We further introduce HAMLET-Aug, which doubles thread length with synthetic on-topic discussion while preserving the thread state. Across 13 models, even frontier models find this setting challenging: the best model achieves 64.0% correctness. We find models can often track discussions to a resolution but tend to over-commit when no resolution exists, while many hallucinated details appear beyond the reference answer. On HAMLET-Aug, doubling thread length shifts hallucination towards additional content. These findings highlight the importance of measuring hallucination in evolving multi-meeting discussions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.