MyoSTAT.AI: An AI-Driven Toolkit for Reproducible Benchmarking of Temporal Segmentation and Shear-Wave Velocity Stabilization on Synthetic Data
Synopsis
This work builds MyoSTAT.AI, a deterministic and fully reproducible benchmarking framework for cardiac ultrasound segmentation and shear-wave elastography (SWE) velocity-field stabilization evaluated entirely on synthetic data, comparing four U-Net-based architectures (2D, 2.5D, 3D, ConvLSTM) through single-variable ablation and six stabilization methods, and reports that the ConvLSTM variant achieved the highest segmentation accuracy (Dice 0.994, IoU 0.987, an 18.6 percentage-point gain over the 2D baseline Dice 0.808), that UNet2.5D with a 3-frame temporal window reached Dice 0.983 at lower latency (433 ms vs. 1193 ms), that TensorRT FP16 deployment sustained 341 FPS on an RTX 3060 and 90.
Figure 1. MyoStat.AI pipeline overview. From cine ultrasound input through preprocessing,
medRxiv · Page 16Interpretation
On the synthetic reference corpora, temporally aware architectures outperformed the single-frame baseline in segmentation: ConvLSTM reached Dice 0.994 and IoU 0.987, UNet3D Dice 0.988 and IoU 0.977, UNet2.5D Dice 0.983 and IoU 0.967, versus the UNet2D baseline at Dice 0.808 and IoU 0.678. Relative to prior work that evaluates segmentation on clinically annotated datasets such as CAMUS and EchoNet-Dynamic, this study places four U-Net-family architectures spanning 2D to temporal designs under one synthetic benchmark and one ablation protocol, reporting Dice, IoU and per-frame latency together. Results come from single runs on synthetic reference corpora (1,200 segmentation frames and 900 SWE velocity-field cases); the authors state that no repeated-run statistics or hypothesis testing were performed, so the ranking should be read as descriptive.
Temporal context helps segmentation, but how it is encoded matters: single-frame input underperformed every multi-frame variant, while UNet2.5D's stacked-channel encoding peaked at a 3-frame window (Dice 0.984, +0.064 over single-frame) and declined monotonically at 5 frames (0.976) and 7 frames (0.956). The work treats temporal window length as its own ablation axis, quantifying the degradation of the stacked-channel representation as the window extends and recording it as failure mode FM-001 (temporal context regression). Based on a single-variable isolation ablation across 17 variants, using a 64-sample synthetic ablation corpus regenerated from global_seed = 42 before each run under a rapid 20-gradient-step mini-train regime.
On synthetic velocity fields, penalized least-squares temporal stabilization (smoothn, lambda in {5, 10, 20}) reduced the temporal CoV of shear-wave speed from 81.07 to 29.37, a 63.8% improvement, with no meaningful difference between lambda values and an edge-preservation score of 1.000 for all three. Against the unstabilized baseline and two SNR-adaptive window methods, this provides a method-to-method comparison among stabilization presets: adaptive_coarse lowered CoV to 6.04 (92.6%) but dropped mean velocity from 3.11 to 0.90 m/s and edge preservation to 0.317, while adaptive_fine raised CoV to 130.67. Evaluated on synthetic velocity fields generated from a fixed random seed with a 30 Hz, 25-frame stabilization window; the authors note these are method-to-method comparisons under controlled conditions rather than absolute predictions of in vivo performance.
Deployment and reproducibility engineering: TensorRT FP16 engines compiled from the same exported ONNX model sustained 341 FPS on the RTX 3060 and 90.4 FPS on the Jetson Orin Nano, and eight independent benchmark runs on the Jetson showed a cross-run CoV of 0.174% with mean inference latency of 11.062 ms (+/-0.019 ms) and p95 below 11.13 ms. The work supplies a benchmark harness with versioned configs, SHA-256 dataset-manifest verification and evidence bundles containing config files, per-variant metrics and checksums, plus a taxonomy of seven engineering failure modes (FM-001 to FM-007) with quantitative evidence and mitigations. All eight runs passed manifest verification; the authors state the procedure measures timing repeatability of an already-trained fixed configuration and does not assess whether training itself or the accuracy metrics are reproducible across initializations.
Perspective
The work is scoped as a computational benchmark in the synthetic domain: the segmentation task is defined as binary target-region mask prediction from synthetic cardiac ultrasound cine frames, and the SWE branch estimates shear-wave speed on synthetic velocity fields, with the excitation source and acquisition physics not modeled. It is meant for engineering and research settings that need reproducible baselines, hardware-specific inference feasibility references and a failure-mode checklist, for example comparing temporal segmentation architectures and stabilization presets on an RTX 3060 workstation and a Jetson Orin Nano embedded platform. The authors state explicitly that this is a synthetic proof-of-concept rather than a validated clinical or operator-independent tool, and that acquisition of real echocardiographic data and clinical validation are required before any claim of diagnostic or deployment readiness.
Training and evaluation masks are produced by the same synthetic generator, so the reported Dice, IoU and CoV values reflect internal consistency with the generator's own geometry rather than independently verified anatomical accuracy; tissue heterogeneity, operator variability and acquisition artifacts intrinsic to real echocardiography remain unassessed. The accuracy ranking comes from single runs, and the loss-function, dropout and augmentation results each come from one 20-gradient-step run per variant with per-variant rather than matched seeding, so a genuine effect cannot yet be distinguished from run-to-run variance, and the related mechanistic explanations (gradient sparsity, weight redundancy, distributional shift) are hypotheses to be tested. Timing repeatability was assessed only on a single serialized engine on a single hardware unit, leaving sensitivity to initialization, hardware or site untested; the manifest checksum currently covers the file listing rather than per-file contents and remains a placeholder until the reference corpora are populated with clinical data. The stabilization window spans 25 frames at 30 Hz, so a causal real-time implementation would additionally incur roughly 0.83 s of window and buffering delay, and that delay rather than millisecond-scale compute overhead governs end-to-end responsiveness. Numerical equivalence between the Python and MATLAB smoothn implementations was assessed only via Pearson correlation on matched outputs, with full MAE and RMSE comparison against a MATLAB reference planned as a validation step prior to clinical deployment.
