Domain-Bootstrapped OCR for Low-Resource Traditional Chinese Painting Inscriptions
Abstract
Traditional Chinese Painting Inscriptions (TCPIs) carry critical information for authorship attribution, artwork dating, provenance tracing, and cultural interpretation, making automatic recognition of such inscriptions an important task for OCR in cultural heritage research. However, TCPIs pose challenges beyond standard document or scene-text OCR due to severe domain shifts, including free-form brush calligraphy, complex background interference, faded ink, and other degradations. These challenges are amplified in low-resource settings, where high-quality annotated TCPI data are scarce and expensive to obtain. To address these issues, we propose a domain-bootstrapped OCR pipeline tailored for low-resource TCPIs. Original paintings are divided into small and large groups and cropped into TCPI patches. The smaller group is used to construct a high-quality seed dataset through multi-model candidate generation and expert verification. A general-purpose OCR model is fine-tuned on this seed dataset to produce an initial TCPI-adapted model, which then generates pseudo-labels for the larger unlabeled group. The same OCR model is subsequently retrained via bootstrapped training on these pseudo-labeled samples using optimized strategies to further enhance recognition performance. Experimental results show that the TCPI-bootstrapped model achieves a clear performance advantage () over existing state-of-the-art methods. Ablation studies further demonstrate that recognition performance consistently improves as the pseudo-labeled dataset expands, highlighting the effectiveness and scalability of the proposed domain-bootstrapping pipeline. The model and benchmark will be released upon acceptance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.