acceptodds
Under review as a conference paper at ICLR 2027

CEPO: Causal Entropy Policy Optimization for Long-Form Narrative Generation

Abstract

Large language models (LLMs) have demonstrated remarkable capabilities in open-ended text generation; however, generating ultra-long novels with long-range causal dependencies, stable character personalities, and coherent narrative arcs remains a substantial challenge. The standard autoregressive generation process is prone to accumulating local inconsistencies, and conventional sequence-level reinforcement learning (e.g., GRPO) suffers from severe credit assignment problems when applied to long-form texts comprising tens of thousands of tokens. To address these challenges, we propose a hierarchical agent-based framework for novel generation that decomposes long-form writing into multiple verifiable stages, utilizing synopses, blueprints, a memory manager, and a Narrative World Graph. Furthermore, to tackle reward sparsity in the optimization of extended texts, we introduce a dense token-level reward mechanism guided by an LLM evaluator. This mechanism translates qualitative narrative feedback into a causal separation of coherent prefixes and contaminated suffixes. Building upon this mechanism, we propose Causal Entropy Policy Optimization (CEPO). By combining suffix-level penalties with an entropy-guided modulation mechanism, CEPO provides a heuristic approximation of token-level contributions around localized failure points, thereby effectively reducing the variance and bias of credit assignment in the generation of ultra-long narratives. Extensive automated and human evaluations on a large-scale dataset of online web novels demonstrate that, despite requiring significantly fewer parameters, the proposed method is competitive with strong baselines on six evaluated dimensions. Consequently, this work establishes an effective paradigm for the high-quality alignment and generation of complex, ultra-long narratives.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.