acceptodds
Under review as a conference paper at ICLR 2027

Learning from Incomplete and Unaligned Lineages for Oracle Bone Script Decipherment

Abstract

Oracle Bone Script (OBS) decipherment seeks the established modern counterpart of an ancient glyph. Related forms preserved across later script periods provide chronologically ordered historical evidence that can help bridge this large visual gap. Yet a cascade of models trained between consecutive stages performs worse than a single model that maps OBS directly to the modern character. We found that the difficulty lies not in the existence of intermediate evidence but in how it is used. For most lineages are incomplete, and the variants of a character at one stage are not paired with those at the next, so each step of the cascade is trained on fragments and restricts cross-period context. We introduce ChronoOBS, which jointly models the six script periods as a chronologically ordered sequence with a shared autoregressive transformer and diffusion head. Missing stages are represented by placeholders, Dynamic Optimal Transport Matching (DOTM) constructs soft cross-period correspondences between unpaired variants, and multi-path decoding aggregates predictions over different subsets of intermediate periods. Under the same underlying architecture, the full ChronoOBS pipeline reaches 24.4% Top-1 accuracy on held-out character categories, compared with 21.8% for direct training under the same architecture and 20.5% for the strongest directly comparable prior image-generation method. On known character categories with period-specific references, the generated intermediate glyphs preserve character identity substantially better than cascaded trajectories. Code and models will be made available.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.