UltraWorld: Learning Interactive Ultrasound World Models from Untracked Clinical Videos with Acoustic Sampling Map
Abstract
World models can support autonomous ultrasound scanning by predicting probe-motion outcomes from local observations. Learning this action–observation relationship typically requires synchronized video–pose pairs, which are costly to collect at scale and generally absent from routine clinical recordings. World models must also account for ultrasound's cross-sectional sampling geometry to follow scanning actions reliably. We present UltraWorld, a self-distillation recipe that transfers clinical video priors into interactive ultrasound world models without real action annotations. Using clinical videos, we adapt a video foundation model into an ultrasound generator conditioned on reference images and masks. The resulting world model predicts from local observations and actions without masks or assets at inference. Our Acoustic Sampling Map (AsMap) encodes probe poses and imaging settings as pixel-wise 3D sampling positions, beam directions, and depths. Experiments show improved prediction fidelity and action following. Across nine simulated closed-loop local planning episodes, UltraWorld reduces mean final distance to the goal and orientation error by 29% and 38%, respectively, compared with visual servoing. Code: https://github.com/sephohi/UltraWorld.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.