acceptodds
Under review as a conference paper at ICLR 2027

EndoWAM: A Grounded World-Action Model for Generalizable Endoscopic Navigation

Abstract

Autonomous endoscopic navigation remains challenging due to tissue deformation, transient occlusions, and rapid viewpoint changes. Existing policies typically predict actions from current observations without explicitly modeling future anatomical dynamics, limiting robustness in safety-critical settings. We present EndoWAM, to our knowledge the first World Action Model (WAM) for generalizable robotic endoscopic navigation. EndoWAM introduces future grounding, which predicts task-relevant target regions in future observations from intermediate denoising features of a video world model. A lightweight diffusion transformer for future target prediction is coupled with a discrete action expert through a shared predictive representation, enabling target-aware dynamics modeling and real-time control with a single denoising pass. We further introduce EndoMotion, a robotic endoscopic motion dataset covering ureteroscopy, esophagoscopy, and endoscopic retrograde cholangiopancreatography (ERCP). EndoWAM consistently outperforms strong baselines and alternative grounding strategies, while exhibiting zero-shot generalization to unseen viewpoints, environments, and targets. These results demonstrate the effectiveness of predictive, target-grounded world-action modeling for robust and generalizable navigation in deformable endoscopic environments.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.