Beyond Pixels: SVG Token Prediction for Historical Document Restoration
Abstract
Historical document restoration (HDR) aims to recover damaged content while preserving the original appearance of historical documents. Diffusion-based and MLLM-based methods have made promising progress in HDR. However, they still provide limited control over character geometry and unintentionally alter preserved character structures. Scalable Vector Graphics (SVG) provide a vector-based representation of character geometry through explicit paths and coordinates. To leverage this representation, we reformulate HDR as **SVG token prediction**, where surviving character structures are retained and only the missing parts are recovered in vector space. To support this formulation, we introduce **VECodex**, the first large-scale vector-based HDR dataset, containing 35.7M aligned restoration instances across four realistic damage types. Its construction combines memory-guided multi-agent quality control with Boolean shape splitting for precise raster–vector alignment. We further propose **Script**, which jointly utilizes pixel-level visual evidence and vector-level character structure to predict missing SVG tokens while preserving observed paths. Experiments on **VECodex** show that **Script** achieves state-of-the-art performance in both character restoration and final document quality, substantially improving Mask IoU and Mask B-F by 23.74% and 38.36% over the strongest baseline, respectively. These results demonstrate the potential of explicit vector-space structure recovery as a new direction for HDR beyond direct pixel generation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.