UltraWorld self-distills an interactive ultrasound world model from untracked clinical videos, cutting terminal distance and orientation error by 29% and 38% versus visual servoing across nine simulated closed-loop planning episodes
Synopsis
UltraWorld presents a self-distillation recipe that first adapts clinical ultrasound videos into a generator conditioned on reference images and anatomical masks, then synthesizes action-video pairs by sampling masks along programmable trajectories through 3D anatomy, distilling the generator into a world model that predicts future observations from local observations and actions alone, and introduces the Acoustic Sampling Map representing probe poses and imaging settings as pixel-wise 3D sampling positions, beam directions, and depths; experiments show improved prediction fidelity and action following, with mean terminal distance and orientation error reduced by 29% and 38% versus visual servoing across nine simulated closed-loop local planning episodes.
Figure 1: Learning ultrasound world models from untracked clinical videos. (a) Clinical videos provide appearance priors, while 3D anatomy provides motion supervision. Together, they yield synthetic action–video pairs for self-distillation. (b) Complex scanning operations change which anatomical cross-sections are imaged and how they are sampled. AsMap encodes the resulting sampling positions, beam directions, and depths to support reliable action control.
arXivInterpretation
A self-distillation recipe for learning action-conditioned ultrasound world models from clinical videos without action annotations: learn a controllable ultrasound video prior conditioned on a reference image and anatomical masks, then sample masks from 3D anatomy along programmable trajectories (sweeps, slide, rock, revisits, changes in imaging depth or field of view) to synthesize video-action pairs paired by construction with their acquisition states, and finally distill the generator into a world model taking local observations and actions as input. Existing ultrasound world models rely on synchronized ultrasound-pose recordings to learn acquisition dynamics, while routine clinical recordings usually lack probe-motion records; this work decouples action supervision from real pairing and constructs it independently from 3D anatomical assets and programmable trajectories, making untracked clinical videos usable. The mask-conditioned video prior is learned on BUV and CAMUS, and anatomy-guided trajectories are constructed from I-SPY1 volumes; all datasets are split at the patient level and derived clips remain in their source patient's partition.
The Acoustic Sampling Map (AsMap), which represents probe poses and imaging settings as pixel-wise 3D sampling positions, unit acoustic beam directions, and sampling depths, expressed in the probe coordinate system of the first frame so that the conditioning changes when imaging depth or field of view changes even with a fixed probe pose. Pose-only conditioning cannot distinguish imaging-scale changes, and camera-style encodings describe viewing rays or projective relationships, whereas a B-mode pixel corresponds to a depth-resolved sample along an acoustic beam; AsMap makes this sampling geometry explicit and is added once before the first transformer block without modifying the internal architecture. Under identical acquisition states, AsMap is best on PSNR 24.994, LPIPS 0.236, latent 0.095, four-frame LPIPS 0.466, and motion error 4.155 against Fourier encoding, MLP encoding, and PRoPE; a parameter-matched MLP (53.49M versus AsMap's 53.48M) remains worse, indicating the gain is not explained by encoder capacity or coordinate convention.
Validation that the mask-conditioned generator preserves realistic ultrasound appearance while following the prescribed anatomical structure, and a demonstration that the world model supports goal-directed local planning in simulation. It connects the generative prior to action-conditioned prediction and further uses it in closed-loop planning with model predictive control and the cross-entropy method, compared against an intensity-based visual-servoing baseline. Generated videos reach SSIM 0.4265 and LPIPS 0.4071, segmentation mean IoU 0.8736, and Dice 0.9260 between generated lesions and conditioning masks (versus 0.9538 between ground-truth segmentations and the same masks); across nine simulated episodes terminal translation and rotation errors are lower than visual servoing.
Consistency evaluation on unseen real clinical frames: patient-held-out BUV and independent BUSI images serve as initial observations, with zero-action and round-trip LPIPS measuring stability and return consistency. It offers a way to probe action sensitivity without paired clinical futures. AsMap attains the lowest zero-action and round-trip LPIPS on both datasets (BUSI 0.042 and 0.062; BUV 0.059 and 0.076); the paper states that round-trip consistency alone does not establish action-following accuracy.
Perspective
The work targets researchers and engineering teams who want to train action-conditioned prediction from routine clinical ultrasound recordings, in settings where 3D anatomical assets are available to construct programmable scanning trajectories and where closed-loop planning can be evaluated in simulation. At inference the method needs neither anatomical masks nor 3D assets, so deployment requires only local observations and acquisition states. The paper states that future work will use a small amount of tracked clinical data for calibration and validation and extend to broader anatomical and device domains and physical robotic systems.
The world model is trained primarily with generator-synthesized action-video trajectories rather than real ultrasound sequences with synchronized probe tracking, so errors or biases in the generative prior may propagate to the distilled world model. AsMap represents acquisition geometry but does not explicitly model tissue-dependent acoustic effects such as scattering, attenuation, and speckle formation, and its metric interpretation assumes calibrated probe poses and imaging geometry. Real clinical ultrasound evaluation is limited to stability and round-trip consistency because paired future observations with known actions are unavailable, and closed-loop robotic planning is currently validated only in simulation and on a limited number of episodes. In addition, several equations and some table values are incomplete in the parsed text, so reproducing exact numbers and hyperparameters requires consulting the original appendices.
