Consistency Checking in Interactive Narratives at Million-Token Scale: A Benchmark and a Dual-Route Evaluator
Abstract
Large language models are increasingly used in text-adventure games and role-playing applications, where narrative histories can span hundreds or thousands of turns. Maintaining factual consistency over such long histories remains challenging; failures to do so can introduce contradictions that disrupt gameplay and reduce immersion. Reliable evaluation is therefore essential, yet the effectiveness of consistency checkers at this scale remains underexplored. To address this gap, we introduce LinConBench (Long Interactive Narrative Consistency Benchmark), to our knowledge the first benchmark for contradiction detection and turn-level localization in interactive narratives extending to thousands of turns. It comprises 2,000 error-bearing narratives derived from 400 human-authored novels across five source-length buckets, ranging from below 64K to 512K–1M tokens. Constructing such a benchmark requires reliable error annotations, yet exhaustively identifying naturally occurring contradictions is labor-intensive and prone to omissions or inconsistent judgments. We therefore adopt a two-stage pipeline that converts novels into multi-turn interactions while preserving their narration, then injects controlled contradictions with annotated error locations and supporting historical evidence. Each sample contains a single injected contradiction expressed in one or more turns and is validated by three reviewers. We further propose DRPV (Dual-Route Proposition Verification), which combines dual-route candidate discovery with proposition consolidation and full-prefix verification. Our experiments with nine baselines show that existing evaluators struggle to localize all erroneous turns, particularly in longer narratives and when historical evidence is more distant, whereas DRPV performs better across multiple metrics.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.