FontBench: A Multidimensional Benchmark for Style-Consistent and Cross-Lingual Font Generation
Abstract
Font generation requires models to jointly preserve character identity, glyph structure, typographic style, and visual quality. However, existing evaluations are often limited to specific scripts, narrow font categories, or reconstruction-oriented metrics, making it difficult to assess the true capabilities of current generative models. We introduce FontBench, a comprehensive benchmark for systematically evaluating font generation across diverse styles, character complexities, and writing systems. FontBench contains 920 fonts organized into three font groups and eight fine-grained subcategories spanning structurally regular to highly stylized fonts. It supports two benchmark tasks, style-conditioned and cross-lingual font generation, together with a multidimensional evaluation protocol covering character readability, structural fidelity, style consistency, and visual quality. Extensive experiments on five specialized font generation models and seven general-purpose image-to-image models reveal complementary strengths and limitations: general models produce highly readable glyphs but often lack precise structural control, whereas specialized font models better preserve font-specific styles yet struggle with complex artistic fonts. Cross-lingual evaluation further reveals directional asymmetry, with English-to-Chinese limited mainly by structural adaptation and Chinese-to-English by style preservation. These results demonstrate the necessity of multidimensional evaluation and establish FontBench as a standardized benchmark for analyzing and advancing font generation systems.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.