What Information Enables Cross-Script Handwriting Layout Synthesis?
Abstract
Handwriting Layout Synthesis generates the spatial arrangement of handwritten text while preserving writer-specific spacing and layout style, providing the structural foundation for handwriting generation. Current systems achieve strong within-script performance but their ability to generalize to unseen scripts is poorly understood. In this work, we present an investigation of the impact of input token representations on this failure. We conduct a controlled empirical study over the state-of-the-art character-compositional representation as well as script-agnostic, explicit glyph-based, and pretrained visual token representations, introducing the latter three to this task. Conducting our study across three distinct scripts representing the morphosyllabic/logographic, abugida and alphabetic writing systems, we observe that neither character-level vocabulary expansion, nor the removal of script-specific information alone is sufficient to achieve robust cross-script generalization. We find that visual representations substantially outperform character-compositional representations on unseen writing systems, while also achieving modest gains in the in-domain setting. Through representation probing and performance comparisons across representations with increasing information, we show that much of the transfer advantage of pretrained visual representations can be attained using coarse geometric information that is linearly recoverable from the learned embeddings, while visual representations retain a modest in-domain performance advantage. Together, these results provide a principled characterization of how representation choice impacts cross-script robustness in handwriting layout synthesis.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.