acceptodds
Under review as a conference paper at ICLR 2027

Masked Alignment Pretraining from Local to Global for Intraoral Image Representation Learning

Abstract

Self-supervised pretraining for intraoral photographs remains underexplored because generic objectives overlook the specific anatomy of the dental arch (\eg bilaterally symmetrical and grouped into six sextants). We present a masked alignment pretraining framework for dental-arch-aware representation learning based on intraoral images. Random, single-tooth, and sextant crops are aligned with dense features at corresponding locations in full-mouth images using geometric overlap as supervision. Masked reconstruction of these crops further captures tooth texture and morphology that complement structural cues. FDI notation (the tooth numbering system) and tooth anatomical relations additionally regularize the tooth representation, so that mirror-symmetric or adjacent teeth, or those in the same sextant are closer, while distant teeth are separated. We also introduce YoungDent, a five-view intraoral dataset with over 5K images from more than 1,000 adolescents captured under non-clinical settings, and combine it with public data into a 10K-scale corpus. Compared with generic SSL baselines, our dental-arch-aware pretraining better captures tooth anatomical organization and improves representation quality, showing clear potential as a foundation for downstream tooth segmentation tasks.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.