acceptodds
Under review as a conference paper at ICLR 2027

Panel2Page: Structured Visual Narrative Generation with Omni-Panel Alignment

Abstract

Structured visual generation requires models to satisfy hierarchical constraints beyond global text-image alignment, including global composition, local semantic correspondence, and cross-panel relational consistency. We study this challenge through text-to-comic page generation, where a model directly synthesizes a complete comic page from a structured prompt specifying layouts, ordered events, recurring characters, textual elements, and cross-panel relations. We propose Panel2Page, a two-stage framework that improves structured correspondence by separating global composition learning from hierarchical semantic alignment while preserving single-pass generation. Page-Level Supervised Fine-Tuning first learns a global structural prior from complete prompt-page pairs, capturing layout organization, reading order, and page-level composition. We then introduce Omni-Panel Alignment (OPA), a training-time strategy that regularizes unified page predictions using detached localized predictions constructed from complementary text-aware, panel-event, and cross-panel views, without introducing additional inference modules. To enable structured supervision, we construct a hierarchical dataset of approximately 20,000 manually verified prompt-page pairs and a benchmark of 320 structured prompts covering diverse layouts and narratives. Extensive experiments demonstrate that Panel2Page achieves the best overall performance among the evaluated full-page and panel-wise generation methods, showing the complementary benefits of global structural learning and hierarchical semantic alignment.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.