acceptodds
Under review as a conference paper at ICLR 2027

OralFlow: A Longitudinal Multimodal Benchmark across Clinical Workflows in Dentistry

Abstract

Multimodal large language models (MLLMs) have expanded dental AI from isolated image interpretation to open-ended clinical interaction. However, existing dental models often frame diagnosis as a static, modality-specific problem. Complete dental workflows require integrating complementary multimodal evidence and continuously retaining, retrieving, and updating evidence across the typical stages of dental care, including diagnosis, treatment planning, and follow-up. To bridge this gap, we introduce OralFlow-Bench, the first benchmark for evaluating longitudinal multimodal dental clinical workflows. It contains 668 cases from two complementary sources: OralFlow-Literature comprises 281 published long-horizon case reports spanning initial presentation, treatment, and follow-up, while OralFlow-Cohort comprises 387 complex, multi-visit hospital cases with 10,326 images across six modalities. We evaluate six MLLMs with six memory methods across five task categories to handle long-context clinical evidence in longitudinal clinical workflows. On end-to-end evaluation across the clinical workflow, models achieve 37.56% diagnosis accuracy and 29.57% treatment and 40.64% follow-up scores, revealing the challenges in handling complete longitudinal dental workflows. Even given ground-truth observations, memory methods yield variation in diagnostic accuracy across MLLMs, ranging from 46.94% to 93.88%, and these differences propagate to subsequent stages, further degrading treatment and follow-up performance. Missing-modality ablations reveal the complementary roles of multiview skeletal model and X-Ray, while noise ablations show that robust performance depends on selective memory updating and retrieval rather than context size alone. Code and benchmark resources are available at https://anonymous.4open.science/r/OralFlow/.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.