HalluciBench: Benchmarking Citation Faithfulness in Large Language Models
Abstract
Large language models are increasingly used to draft scholarly text and assemble bibliographies, yet citation hallucinations have become a visible failure mode in the ML publication pipeline, with recent venue guidance and community scrutiny around NeurIPS, ICLR, and AAAI highlighting how incorrect or unverifiable references can propagate into submissions. We introduce HalluciBench, a benchmark for citation faithfulness where models must output a single BibTeX entry or abstain with REFUSE. HalluciBench contains a dataset of 6,000 records and five prompt variants ranging from exact-title requests to progressively noisier partial-metadata cues, enabling controlled analysis of when models retrieve faithfully versus guess. We evaluate outputs using a deterministic, field-level protocol covering parseability, author accuracy, title fidelity, venue, year, and DOI. Across evaluated models, citation formatting and factual correctness remain unreliable: BibTeX parseability ranges from 0.8 % to 99.8 %, year exact-match reaches only 0.6 % to 79.7 %, and author-set F1 varies substantially. Models also differ sharply in calibration, with correct refusal rates on no-answer prompts ranging from 1.0 % to 100 %. HalluciBench provides a transparent, reproducible platform for tracking progress toward verifiable citation generation and appropriately calibrated abstention.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.