Document Post-Train for Image Generation
Abstract
Document page generation with text-to-image models requires accurate rendering of long passages and precise layout control, enabling automated creation of text-rich materials. However, existing models often suffer from local text defects and inaccurate layouts. To address these limitations, we introduce a post-training method covering both data and optimization. On the data side, we construct an automated pipeline turning real PDF pages into OCR-grounded dense captions with layout and visual features. Regarding optimization, we design legibility rewards based on VLM and region-level credit assignment specifically to address local text defects. Evaluations demonstrate that an open-source model trained with our method achieves commercial-level performance on our benchmark, ranking second only to GPT Image 2 overall, and this improvement also generalizes to other text-rich tasks. Our further analysis reveals the ineffectiveness of image-level optimization and OCR rewards, and shows that reinforcement learning reduces sampling needs while retaining diversity for test-time scaling. We believe these findings offer a practical path toward high-quality document generation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.