Soft Temporal Scoring Using a Foundation Model: Optimal Frame Selection for Improved ONSD Measurement in Ultrasound Videos
Synopsis
The study presents a sparsely supervised AI framework in which frozen ultrasound foundation model (USFM, ViT-B) embeddings of 768 dimensions per frame are passed to a lightweight bidirectional LSTM temporal head that outputs a 0–1 frame-quality score, trained with Gaussian soft labels that peak at expert-marked key frames and decay smoothly with frame distance (width σ = 5 frames); in subject-level five-fold cross-validation on 18 subjects and 323 ultrasound videos spanning nine controlled acquisition sweep types (about 45,500 frames), it selected a usable frame in 82.2% of sweeps containing key frames with a mean minimum distance of 3.07 frames, exceeding the strongest training-free baseline (USFM feature cosine similarity, 49.2%) and a hard-label model (71.9%).
Fig. 1. ONSD frame-quality scoring framework. A frozen USFM encoder and temporal head produce per-frame scores. Gaussian soft targets from labeled key frames supervise the temporal head during training, while inference selects the maximum-scoring frame.
medRxiv · Page 2Interpretation
ONSD key frame selection is formulated as sparsely supervised per-frame quality scoring with Gaussian soft labels: each annotated key frame induces a Gaussian curve and each per-frame label is the maximum over all keys, giving a value of 1 at the key and smooth decay with temporal distance. Prior approaches relied on anatomical segmentation, hand-designed echogenicity templates, or hard 0/1 per-frame classification; this work replaces the binary boundary with continuous graded targets, which the authors argue better reflects the similarity of neighbouring ultrasound frames and avoids forcing abrupt decisions at valid-region boundaries. Evidence comes from the design and the comparison experiments: the soft-label model reached 82.2% H@1 versus 71.9% for the hard-label model, and in the single-sweep example of Figure 3 the boundary key frame (frame 81) scored 0.30 under the hard-label model and 0.87 under the soft-label model.
On top of frozen USFM features, only a lightweight bidirectional LSTM temporal head (hidden size 128 per direction, output 256) with a linear projection and sigmoid is trained to emit one 0–1 score per frame, and inference selects the maximum-scoring frame. Combining foundation-model features with temporal scoring that needs only sparse frame labels: freezing the encoder limits task-specific learning to few parameters, which the authors state can drastically reduce the amount of ONSD data needed, while the temporal head scores each frame in the context of the whole sweep rather than independently. Ablation showed that removing the LSTM lowered H@1 to 74.4% and raised MMD to 3.76, indicating the sequence model contributes beyond per-frame features; training used masked binary cross-entropy with Adam for 200 epochs, learning rate 5×10⁻⁵, batch size 8.
Subject-level five-fold cross-validation over 18 subjects and 323 sweeps across nine controlled sweep types yielded H@1 82.2%, H@3 84.3%, H@5 86.0%, and MMD 3.07 frames for the soft-label model. Baselines include random selection, RMSE, NCC, USFM feature cosine similarity (the strongest training-free baseline at 49.2%), and hard-label models with frozen or fine-tuned encoders, allowing the gains to be attributed to soft labels and temporal modelling rather than to foundation-model features alone. The advantage is consistent across metrics (H@5 86.0% versus 78.9% for the hard-label model), but the evaluation is a small single-site cohort and no separate test set was used.
Broken down by acquisition view, the soft-label model matched or exceeded the hard-label model in H@1 across all nine views, with the largest gains for leftward translation (58.3% to 100.0%) and upward fanning (41.7% to 61.1%). This indicates that graded supervision holds across several probe motion patterns rather than only one acquisition type; MMD improved for most views but increased for upward translation (1.29 to 2.43 frames) and downward fanning (3.20 to 9.00 frames). Based on the per-view statistics in Table 2, where the sample size per view is small.
Perspective
The framework applies to settings where a usable ONSD frame must be picked out of an ultrasound sweep, aimed at users who want ONSD assessment in low-resource or austere settings with personnel of limited ultrasound experience; the authors position it as a first step towards automated acquisition guidance, automated quality assessment, and objective ONSD measurement. It requires only sparse frame-level labels rather than segmentation masks, and inference can run on CPU (3.2 ± 0.3 s per sweep on an NVIDIA Tesla V100 GPU and 40.1 ± 2.4 s per sweep on an Intel Xeon Gold 6130 CPU node), although the authors note further optimization may be required for real-time use. As reported, the result covers locating a valid frame across nine controlled acquisition sweep types, and planned next steps are multi-site generalization and integration of frame selection with automated ONSD measurement.
Inter-annotator agreement was not evaluated, so σ was not calibrated to annotator variability, although a post-hoc sensitivity analysis showed stable performance across σ of 3, 5, 7, and 9 (for example, H@1 81.8 ± 1.9% at σ = 3 and 82.2 ± 2.3% at σ = 5). Localization metrics exclude the 80 sweeps (24.8%) that contain no key frame, and whether such sweeps can be rejected at inference was not evaluated. The evaluation rests on a single-site cohort of 18 subjects, leaving generalizability across subjects, scanners, operators, and sites an open question, and no separate test set was used. In this small cohort, end-to-end fine-tuning underperformed frozen-feature training, and the cause was not investigated. The authors note that H@1 and MMD are more discriminating metrics here because learned methods tend to rank adjacent frames together. The work evaluates frame selection only, and whether downstream ONSD measurement accuracy or clinical decision-making improves is not established. The text is also a medRxiv preprint that has not been certified by peer review, so the reported numbers await independent reproduction.
