acceptodds
Under review as a conference paper at ICLR 2027

From Signs to Sequences: Compositional Synthesis for Cuneiform Recognition Across Periods and Collections

Abstract

Training data that pair cuneiform line images with sign sequences are scarce, even though isolated sign images and machine-readable texts are available separately. This scarcity keeps large collections of digitized tablets difficult to search and analyze automatically. We introduce a compositional supervision method that combines these independently available resources. Given a normalized sign sequence from a text corpus, we sample corresponding real sign crops from a tablet-disjoint visual dictionary (augmented for rare signs) and compose them into a synthetic line image. These image–sequence pairs pre-train three recognition architectures—a CRNN, an SVTRv2-inspired CTC model, and PaddleOCR-VL—before adaptation on a limited number of annotated tablets. In the principal evaluation, sign locations are supplied by ground-truth annotations and the input lines are reconstructed from their crops. With a fixed sign inventory, adaptation on 100 Ur III tablets yields sign error rates (SER) of 18–25% on in-period data. Removing compositional pre-training increases error by 24–52 percentage points across architectures and evaluation domains. Controlled experiments within the SVTRv2-inspired model further show that multi-sign sequence training provides substantially greater benefits than isolated-sign training, whereas preserving attested sign order and co-occurrence produces smaller, domain-dependent gains. Adaptation primarily improves recognition within the target period, while cross-period and cross-collection transfer remains difficult. An exploratory modular pipeline with predicted sign locations and reading order reaches 38.05% whole-side SER on reused in-period tablets, versus 68.92% for its strongest original geometric baseline. These results establish a practical way to repurpose existing Assyriological image and text collections as sequence-level supervision, while clarifying the remaining challenges for scalable cuneiform image-to-text systems.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.