acceptodds
Under review as a conference paper at ICLR 2027

Overjet-1: A Multi-modal Generalist Vision-Language Model for Grounded Dental Image Understanding

Abstract

A complete dental examination is an inherently multi-modal, multi-view diagnostic task: no single imaging source provides complete insight into underlying pathologies. Clinicians aggregate evidence across sources such as panoramic and intraoral radiographs, color photographs, cephalograms, periodontal charts, and 3D CBCT scans. With the advent of deep learning and vision-language models (VLMs), significant progress has been made in dental image analysis. However, current VLMs and task-specific models process single images in isolation, missing crucial cross-modal context and allowing errors to compound across findings. We present Overjet-1, a generalist dental VLM built on Gemma 4 26B-A4B that natively processes multi-image, multi-modal inputs. Overjet-1 is post-trained on 1,631,594 supervised records spanning 122 dental tasks. These records are organized through the Dental Graph, a case-level representation that unifies tooth segmentations, finding regions, clinical measurements, and 3D geometry across all available multi-modal sources. The graph supplies the spatial coordinates for each training target, links observations of the same tooth across views, and constructs chain-of-thought (CoT) reasoning paths through graph traversal. We introduce Dental Consistency Reinforcement Learning (DC-RL), a novel framework extending Group Relative Policy Optimization (GRPO). DC-RL computes verifiable, deterministic reward signals directly from grounded facts in the Dental Graph, simultaneously optimizing diagnostic correctness and spatial localization without relying on extensive manual annotation or CoT distillation. Evaluated on MMOral-OPG-Bench, Overjet-1 achieves a new state of the art of 71.5%, compared to the best baseline of 61.3%. We also introduce Overjet-Bench, a more comprehensive benchmark covering 9 task families across single-image, multi-image, and multi-modal setups. Overjet-1 achieves a macro score of 78.8% on Overjet-Bench, outperforming top proprietary and open-source baselines (33.5%) significantly on real-world tasks. To our knowledge, Overjet-1 is the first model natively trained and evaluated on both multi-image and multi-modal dental tasks with grounded supervision. Code is available at https://anonymous.4open.science/r/Overjet-1-code-D205; Overjet-Bench will be released upon acceptance.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.