EndoWAM: A Grounded World-Action Model for Generalizable Endoscopic Navigation
Abstract
Autonomous endoscopic navigation remains challenging due to tissue deformation, transient occlusions, and rapid viewpoint changes. Existing policies typically predict actions from current observations without explicitly modeling future anatomical dynamics, limiting robustness in safety-critical settings. We present EndoWAM, to our knowledge the first World Action Model (WAM) for generalizable robotic endoscopic navigation. EndoWAM introduces future grounding, which predicts task-relevant target regions in future observations from intermediate denoising features of a video world model. A lightweight diffusion transformer for future target prediction is coupled with a discrete action expert through a shared predictive representation, enabling target-aware dynamics modeling and real-time control with a single denoising pass. We further introduce EndoMotion, a robotic endoscopic motion dataset covering ureteroscopy, esophagoscopy, and endoscopic retrograde cholangiopancreatography (ERCP). EndoWAM consistently outperforms strong baselines and alternative grounding strategies, while exhibiting zero-shot generalization to unseen viewpoints, environments, and targets. These results demonstrate the effectiveness of predictive, target-grounded world-action modeling for robust and generalizable navigation in deformable endoscopic environments.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.