acceptodds
Under review as a conference paper at ICLR 2027

TabletVLM: End-to-End Cuneiform Tablet Transcription with a Spatially Grounded Vision-Language Model

Abstract

Transcription of cuneiform tablets requires solving spatial localization, reading-order reconstruction, and sign recognition across substantial variations in image quality, tablet condition, and scribal style. Existing approaches address these challenges through discrete pipeline stages, requiring different datasets, evaluations, and models, or attempt direct end-to-end training. These difficulties are compounded by the low-resource nature of cuneiform and by existing datasets that are often incomplete, inconsistently annotated, or derived from heterogeneous sources. We propose a single vision-language model that directly transcribes cuneiform tablets without explicit sign detection supervision or reading-order reconstruction and investigate how a training curriculum with supervision at different visual and sequential scales can be used to improve transcription performance. Among the four curricula tested, the model with the best trade-off between transcription and localization performance achieves a mean tablet-side sign-error rate of 27.24% on 4,805 held-out tablet sides, while its peak cross-attention falls within the corresponding ground-truth sign region for 78.27% of 879 held-out annotated signs. More broadly, our results suggest that end-to-end VLMs can learn both accurate sequence prediction and spatial grounding in low-resource OCR settings with limited and uneven supervision.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.