Skip to main content
Back to timeline
arXivSource publication:

A low-rank constraint turns the ultrasound video latent space into a readable cardiac-cycle trajectory, giving ED/ES frames without extra training

Synopsis

The work introduces LRM-Functa, which imposes a low-rank constraint on the time-resolved modulation vectors of VidFuncta (mt = v + Bβϕt), so that the latent space of cardiac ultrasound videos forms periodic spiral trajectories; this allows direct readout of end-diastolic (ED) and end-systolic (ES) frames without additional model training, keeps stable reconstruction and ejection-fraction prediction at the extremely low rank k = 2 on EchoNet-Dynamic, and generalizes to a POCUS cardiac view and to lung-ultrasound B-line classification.

Source-provided article image: Low-Rank-Modulated Functa: Exploring the Latent Space of Implicit Neural Representations for Interpretable Ultrasound Video Analysis

Interpretation

The authors first inspect pairwise cosine similarities of VidFuncta's frame-wise modulation vectors ϕ and find a low-rank pattern with clear periodicity along the side diagonals; based on this they constrain each frame-specific modulation vector to mt = v + Bβϕt, where Bβ is a learnable rank-k subspace. Earlier Functa-family methods (Functa, VidFuncta, Latent-INR and others) focused on reconstruction and downstream tasks, leaving latent-space structure and interpretability largely unexplored; this work is, to the authors' knowledge, the first to analyze and structure the latent space of Functa-based frameworks. Grounded in a cosine-similarity visualization of the VidFuncta latent space, with training and evaluation on EchoNet-Dynamic (10,030 four-chamber cardiac videos) and comparison against COIN++, MedFuncta, VidFuncta and a convolutional autoencoder.

The low-rank latent space is directly visualizable at k = 2: LRM-Functao forms a periodic spiral trajectory, whereas LRM-Functab collapses the trajectory to a line; projecting ϕ onto the PCA principal motion direction and applying a Savitzky–Golay filter yields a signal whose peaks and valleys give unsupervised ED and ES indices. ED/ES identification no longer needs an additional trained model but is read directly from the latent trajectory; the authors describe this as the first structured, interpretable latent representation in an INR-based architecture. On the EchoNet-Dynamic test set at k = 2, LRM-Functab reaches ED MAE 2.26 ± 2.5 and ES MAE 1.98 ± 2.1, close to the fully supervised approach the authors cite (ED MAE 2.1, ES MAE 1.7); at k = 512 LRM-Functao reaches ES MAE 2.16 ± 2.2.

Reconstruction stays stable at very low rank: at k = 2 and k = 4 LRM-Functa is the only method maintaining stable reconstruction quality, and at higher ranks LRM-Functao matches or exceeds all compared methods. Where prior methods degrade at low rank, this work compresses each frame to k = 2 (described as roughly 3000× compression) while remaining usable, with k = 512 corresponding to roughly 25× compression. Evaluated with PSNR and SSIM3D on the predefined EchoNet-Dynamic test set of 1277 samples, plotted as a function of k (Figure 5).

Compressed reconstructions retain clinically relevant information: at k = 2 LRM-Functao gives ejection-fraction MAE 5.29 ± 0.1, RMSE 7.14 ± 0.1 and R² 0.67 ± 0.01, close to the original-video upper bound (MAE 4.93 ± 0.3, RMSE 6.71 ± 0.4, R² 0.70 ± 0.03); lung-ultrasound B-line classification at k = 512 reaches AUROC 91.7 ± 3.7 versus 91.5 ± 5.9 on original videos. It links high compression rates directly to preserved downstream clinical performance, and shows Functa-based methods retain diagnostically relevant features for B-lines while the convolutional autoencoder fails to retain the needed detail. Downstream models are ResNet 18-3D trained for 20 epochs with 5-fold cross-validation on reconstructed videos; ejection fraction uses MAE/RMSE/R² and B-lines use accuracy/F1/AUROC; the lung-ultrasound data are 200 adult emergency-department clips with 40 cases held out for testing.

Perspective

The results target compact representations and unsupervised cardiac-phase readout for ultrasound videos (mainly four-chamber cardiac and two-chamber POCUS views, plus the lung-ultrasound B-line task); the intended setting is one where ED/ES indices are read from the latent trajectory without training a separate phase-detection model, and where each frame is compressed to between k = 2 and k = 512 coefficients. For researchers who want to reuse the representation for anomaly detection or congenital heart disease classification, the authors list both as future work.

The out-of-distribution POCUS evaluation uses only 17 videos and the lung-ultrasound test set has 40 cases, so sample sizes are limited and stability across centers and devices remains an open question. The authors note that MAE on POCUS is lower than on EchoNet-Dynamic and suggest this is likely due to the lower frame rate, an explanation that still awaits confirmation. They also note that B-line detection needs finer spatial detail, so that task is evaluated only at k = 512, leaving behavior at very low rank unreported. In the conclusion they propose improving reconstruction fidelity to preserve fine structural detail needed for assessing left ventricular hypertrophy, indicating current fidelity may not yet suffice for such fine-structure tasks. This is a full-text read, but the numeric curves in Figure 5 and the trajectory details in Figures 3 and 4 are presented as images without point-by-point values in the text, so the exact magnitude of reconstruction quality versus k can only be judged from the prose.

Sources