A video-to-model framework uses U-Net and a spatio-temporal CNN to identify CBF–CLF–QP suture-thread parameters from video, reproducing thread behavior with low tracking error on unseen motions
Synopsis
The work presents a video-to-model framework that first extracts an ordered suture-thread node trajectory from a monocular video using U-Net segmentation and curve tracking, then uses a spatio-temporal CNN to estimate the connectivity, stiffness, and velocity-retention parameters of a structured CBF–CLF–QP model, which simulates the thread under a user-defined needle velocity; on unseen virtual materials and a non-training flower-petal needle motion, the configured model reproduces the observed thread behavior with low tracking error and reduces manual parameter tuning.
Figure 1: Overview of the proposed video-to-model framework for automatic thread-model configuration. An input video is fed into a perception module to detect the centerline of the DLO. This object is then fed into a parameter estimator module which enables a CBF–CLF–QP simulation of the object.
arXivInterpretation
An end-to-end video-to-model pipeline automatically converts observed suture-thread motion into a parameter vector for the CBF–CLF–QP model, instantiating a thread-specific model without manual tuning. Previously the CBF–CLF–QP parameters had to be selected so that the simulated thread matched the observed response to the same needle motion; this work instead identifies parameters directly from video observations. Evaluated on virtual materials not in the training set and on an asymmetric flower-petal needle motion absent from the training motions, with reported parameter estimates and QP modeling errors plus trajectory reconstruction comparisons.
The perception module uses a four-level U-Net to produce a thread probability map, then temporal filtering, morphological closing, skeletonization, and graph traversal (using the previous frame's direction at self-crossings, with Euler-path fallback) to yield a temporally consistent ordered node trajectory. It combines segmentation, post-processing, and curve tracking into a trajectory compatible with the CBF–CLF–QP discrete node representation, handling self-crossings and endpoint-direction consistency. The segmentation dataset is rendered from simulated trajectories: ten videos of 120 frames each, six for training, two for validation, two for testing; the loss combines weighted BCE, Dice, and clDice; ordered centerlines were obtained in most frames with a reported trace success rate.
The spatio-temporal CNN takes a twelve-dimensional per-node per-frame feature vector (needle-relative position, normalized velocities, node spacing, local bending, normalized CBF and CLF, position along the thread, needle velocity, speed context) and outputs three normalized parameters through a node encoder, 1-D convolutions along the node dimension, spatial attention, a GRU, and four temporal attention heads. It casts parameter identification as a learned mapping and adds a cross-motion consistency loss so the network identifies material-dependent parameters rather than the particular excitation motion. Training data comprise 200 virtual materials, ten informative needle motions each, 400 time steps per trajectory, split by material into 140 training, 30 validation, and 30 test materials, trained with a supervised parameter loss and a cross-motion consistency loss.
Beyond the parameter-level loss, a differentiable multi-step QP fine-tuning stage computes an average tracking loss over a 12-step rollout generated by the predicted parameters and backpropagates it, so the predicted gains yield low tracking error in simulation. It notes that a parameter-level loss alone does not ensure low simulation tracking error and therefore adds trajectory-level fine-tuning; this loss was moved to the fine-tuning stage because it was computationally and memory intensive and produced unstable gradients during training. Fine-tuning used 48 randomly selected rollout windows per epoch, batch size two, and a learning rate; because the CBF–CLF–QP simulator is differentiable, the trajectory error propagates through all 12 simulation steps.
Perspective
The result targets planar motion, monocular video, and a thread pulled by a needle in a fluid, and assumes the material and its model parameters stay constant throughout the video; parameter identification averages window-level estimates in normalized parameter space over multiple overlapping temporal windows of at most 400 frames, starting 200 frames apart, with every second frame fed to the network. For a user, this means that given an observable needle–thread interaction video and a chosen CBF–CLF–QP structured model, one can automatically obtain a thread-specific parameter set and use it directly for simulation without tuning each parameter by hand; the authors also state that future work will address more robust visual tracking, parameter-sensitive excitation trajectories, and experiments with real suture threads made from different materials.
The results rest on simulated rendered virtual materials and virtual needle motions; behavior on real suture threads, real illumination, and occlusion remains to be tested. The authors note that occasional segmentation and centerline-tracing errors introduce biased node observations that affect parameter predictions, and that some excitation trajectories do not sufficiently distinguish individual parameter effects, making unique identification difficult. In addition, several numerical values in the loaded text (such as the time step, frame rate, total duration, trace success rate, parameter ranges, and some hyperparameters) are missing at equations and tables, so their exact values cannot be confirmed here; readers needing precise reproduction should consult the original figures and tables.
