Skip to main content
Back to timeline
arXivSource publication:

Robotic ultrasound drives real-time CBCT slice updating: USCorUNet cuts forward-backward residual by about 53% and updates a slice in 11.25 ms

Synopsis

The work proposes a deformation-aware CBCT updating framework that uses robotic ultrasound as a dynamic proxy to infer tissue motion: after hand-eye calibration initialization and LC2-based rigid refinement, a lightweight network, USCorUNet, estimates dense bidirectional deformation fields from adjacent ultrasound frames, which are spatially regularized and transferred to the CBCT reference slice, enabling real-time end-to-end CBCT slice updating without additional radiation exposure, validated on phantom and in vivo data for both deformation estimation and ultrasound-guided CBCT updating.

Source-provided article image: Robotic Ultrasound Makes CBCT Alive

Interpretation

A workflow-compatible CBCT-ultrasound pipeline that chains calibration, LC2 rigid refinement, ultrasound deformation estimation, and ultrasound-informed CBCT slice updating into an end-to-end system. Prior work using electromagnetic and optical tracking mainly achieved rigid CBCT-ultrasound alignment, leaving deformation-aware multimodal updating an open challenge; this work integrates rigid alignment with deformation transfer into a complete real-time chain. The method section specifies the full module structure and formulas (calibration transform, LC2 local linear approximation, Gaussian probe compression profile, Euclidean-distance-transform-based spatial weighting), validated on four datasets: in vivo forearm/upper-arm, a pork-tissue gel phantom, a chicken/pork gel phantom, and a Kyoto Kagaku US-22 abdominal phantom.

USCorUNet: a ResUNet-style context encoder-decoder plus a shared-weight correlation encoder builds local correlation volumes from a five-channel input (I0, I1, their difference, and both gradient magnitudes), decoding dense bidirectional fields at 1/8 resolution. Compared with directly using a general optical flow model such as RAFT, the network explicitly introduces correlation-based matching and is trained with a combined objective of optical flow distillation, confidence-weighted photometric consistency, and regularization (edge-aware smoothness plus a Jacobian-based folding penalty), with weights λflow=1, λphoto=0.2, λreg=0.05. On Dataset A, USCorUNet matches RAFT in photometric alignment (MAE 0.05, NCC 0.85), slightly improves bone-mask Dice (90.62% vs 90.31%), reduces the mean forward-backward residual from 1.81 to 0.85 (about 53%), and lowers the folding ratio from 0.24% to 0.13%; ablations show removing the correlation branch drops NCC to 0.61 and removing photometric supervision drops Dice to 85.45%.

Under the force-stratified protocol against DefCor-Net, USCorUNet is more stable at higher force levels: Dice 87.5% vs 82.6% at 6 N and 89.8% vs 87.8% at 5 N, with smaller standard deviations. The gains appear mainly under larger, more challenging compressions, indicating better adaptation to large probe-induced deformations. Evaluated on the same in vivo ultrasound dataset (Dataset A) under DefCor-Net's force-stratified protocol, reporting bone-mask Dice for the I1→I0 direction.

In ultrasound-guided CBCT updating, the method achieves a quality-efficiency trade-off: on Dataset D, MAE 0.33, SSIM 0.22, Dice 82.22%, and 11.25 ms end-to-end, about 5× faster than RAFT (56.24 ms) and about 512× faster than LC2-FFD (5764.26 ms). It slightly outperforms RAFT while avoiding tearing artifacts, and clearly exceeds LC2-FFD in Dice (82.22% vs 58.91%) without geometric distortion. All three methods were evaluated under identical conditions with end-to-end runtime; the authors note that absolute metrics are affected by robotic artifacts, so this serves as a controlled relative comparison.

Perspective

The framework targets robotic ultrasound-assisted interventions, suited to operating rooms already equipped with robotic ultrasound and CBCT and able to perform hand-eye calibration; its design goal is to keep static CBCT deformation-consistent under probe-induced and externally induced motion rather than to replace CBCT imaging itself. Methodologically, deformation is estimated from adjacent ultrasound frames, made physically more plausible through optical flow distillation and regularization, then corrected with a Gaussian probe compression profile and transferred to the CBCT region of interest via Euclidean-distance-transform weighting, so its scope relates to locally ultrasound-visible soft tissue. The authors note that future work could bring semantic segmentation into the registration pipeline to further refine deformation details in complex anatomical regions.

CBCT updating experiments are mainly on an abdominal phantom (Dataset D), while in vivo data support deformation estimation and force-stratified evaluation, so behavior in real patient anatomy and clinical workflow still needs further observation. The authors explicitly state that absolute metrics are affected by robotic artifacts, so Table 4 should be read as a controlled relative comparison rather than an absolute accuracy benchmark. Deformation estimation relies on pseudo-labels distilled from RAFT, so its quality ceiling depends on the teacher model; training data scale and acquisition conditions (KUKA LBR iiwa 14 R820, Siemens ACUSON Juniper 5C1 probe, Loop-X Imaging Ring) are given in the text, but reproducibility across centers remains to be tested. In addition, this is a fast-parse version in which the visual details of Figures 1-4 are not expanded in the text, so assessing the specific appearance of artifacts and tearing would require consulting the original figures.

Sources