acceptodds
Under review as a conference paper at ICLR 2027

NovelAudit: A Benchmark for Open-Ended Narrative Consistency Auditing in Full-Length Novels

Abstract

AI-generated stories and scripts create a practical need for narrative consistency auditing. We begin with complete human-authored novels, where apparent contradictions may have distant explanations. Existing tasks emphasize supplied claims, predefined error categories, or detected errors, and largely omit explained reversals from required discovery outputs. We introduce NovelAudit, which asks agents to discover explained reversals and unresolved inconsistencies without supplied cases and support them with role-structured quotations. We construct references for 13 principal and 11 extended novels; evaluation separates task completion, case discovery, and explanatory sufficiency. On the principal novels, we compare DeepSeek across six runtimes, with supplementary Qwen and GPT–Codex conditions. Differences in runtime and model–runtime adaptation are clearest in completion rate, while single attempts discover limited, changing case subsets. Complementary discoveries across rounds expand coverage at increasing input cost per new match. Human audits link Explanation-Path F1 to explanatory sufficiency and identify valid predicted cases missed by matching, distinguishing discovery from evidence quality.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.