SAW conditions surgical video diffusion on four lightweight signals, cutting CD-FVD to 199.19 and lifting rare-action F1 from 20.93% to 43.14%
Synopsis
The work proposes Surgical Action World (SAW), which reformulates video-to-video diffusion as trajectory-conditioned surgical action synthesis conditioned only on four lightweight signals—a language prompt, a reference first frame, a tissue affordance mask, and 2D tool-tip trajectories—fine-tunes LTX-Video on a custom-curated dataset of 12,044 laparoscopic clips with a depth consistency loss, reaches CD-FVD 199.19 (vs. 546.82 for SurgSora) and FVD 224.28 on held-out test data, and demonstrates that augmenting rare actions with generated videos improves action recognition on real test data (clipping F1 20.93% to 43.14%; cutting 0.00% to 8.33%) plus initial feasibility of rendering tool-tissue interaction videos from simulator-derived trajectories.
Interpretation
A curated dataset of 12,044 laparoscopic video clips annotated with video-level tool-action labels (clipping, grasping, cutting, dissecting), active tool class (grasper, hook, clipper, scissors), tissue affordance, and frame-level tool-tip pixel coordinates. Prior surgical video generation methods relied on per-frame segmentation masks or spatio-temporal scene graphs that are expensive or hard to obtain at inference; this dataset compresses supervision into lightweight spatiotemporal annotations. Clips come from 101 source videos (21 YouTube videos plus the public HeiChole, Cholec80, SurgVU, and CRCD datasets), split by source video into 11,502 training and 542 test clips, standardized to 81 frames at 25 fps and 1024x576.
A conditional video diffusion approach built on LTX-Video and fine-tuned with IC-LoRA that reformulates video-to-video diffusion into trajectory-conditioned surgical action synthesis, requiring only a language prompt, reference first frame, tissue affordance mask, and 2D tool-tip trajectory at inference. Compared with HieraSurg's dense annotations, SG2VID's structured scene graphs, and SurgSora's limited inference window (W = 21), the formulation avoids expensive annotations or structured intermediates at inference time. Fine-tuned on a single NVIDIA A100 for 7,500 steps with IC-LoRA (alpha = 128, lr = 2x10^-4, AdamW, bfloat16), with 50 denoising steps and guidance scale 3.5 at inference, generating 81-frame videos.
A depth consistency loss (LDC): during training, depth-mask videos are produced with Depth Anything V2, and a cross-attention layer plus projection head reconstruct masked depth latents from denoised RGB latent tokens under a Smooth L1 loss, enforcing geometric plausibility in the Z dimension without depth inputs at inference. All spatial conditioning signals are 2D inputs, so the loss places depth supervision only in training rather than adding a depth condition at inference. Ablation shows removing LDC raises CD-FVD from 199.19 to 207.59, which the authors read as reduced temporal consistency of tool and tissue movement; the full model is described as the most balanced across metrics.
On held-out test data SAW achieves the lowest FVD (224.28) and CD-FVD (199.19), ahead of WAN (439.60, 429.67), LTXb (319.37, 504.31), and SurgSora (541.61, 546.82), and demonstrates two downstream uses: augmentation with 287 synthetic videos (110 clipping, 177 cutting) raises spatiotemporal CNN F1 on real test data from 20.93% to 43.14% for clipping and 0.00% to 8.33% for cutting; and simulator-derived instrument segmentations, tip trajectories, and tissue affordance from an Isaac Lab simulator drive generation of tool-tissue interaction videos. Generation quality and downstream task benefit are reported within one framework rather than stopping at visual metrics. Generation metrics are reported on a held-out test set; augmentation is evaluated on real test data with two recognition models (spatiotemporal CNN and ViT), where ViT cutting rises from 30.77% to 47.06% but clipping slips from 84.62% to 83.64% and grasping slips slightly in both; the simulation part is explicitly called an initial proof-of-concept.
Perspective
The results target laparoscopic surgical video generation, particularly tool-tissue interaction in cholecystectomy scenes, and are meant for researchers and simulation developers who need controllable synthesis of surgical action videos. The approach removes the need for per-frame segmentation masks or scene graphs at inference, requiring only a language prompt, reference first frame, tissue affordance mask, and 2D tip trajectory, which makes it easier to scale; it also provides a reusable pipeline for rare-action augmentation and for rendering interaction videos from simulator-derived trajectories. The authors position the simulation part as an initial proof-of-concept and note that future work must address instrument joint kinematics, more instruments and scenes, and real-time inference.
Ablations show that removing affordance or language slightly improves some metrics (w/o Affordance FVD 223.78; w/o Language FVD 219.23) while the authors keep all conditioning signals for overall balance, so when each signal genuinely matters remains an open question. Augmentation gains are not uniform across the two recognition models: the spatiotemporal CNN improves clearly on clipping and cutting, whereas ViT slips slightly on clipping and grasping, indicating that synthetic-data benefit depends on model and action class. The simulation part is qualitative with no reported quantitative metrics, and real-time behavior and instrument joint kinematics are not yet validated. Generated videos are fixed at 81 frames, leaving longer-horizon behavior and performance across more instruments and anatomical scenes to be examined.
