VA-Adapter lets an ultrasound foundation model guide the echo probe, cutting trained parameters about 33-fold while lowering guidance error below strong baselines
Synopsis
The work proposes a Vision-Action Adapter (VA-Adapter) inserted into the deep layers of a frozen ultrasound foundation model image encoder (EchoCLIP, USFM, BiomedCLIP) to online-inject understanding of individual 3D cardiac structure by encoding historical vision-action sequences; on a dataset of 178 adults, 356 expert scans and 1.31M image-action pairs, it reaches lower translation and rotation mean absolute error than strong probe-guidance baselines with roughly 2.61M-3.97M trainable parameters.
Interpretation
It introduces VA-Adapter, which models probe guidance as vision-action sequence interaction embedded in the deep layers of a foundation model image encoder, letting the model infer individual cardiac anatomy from historical frames and relative poses. Prior probe-guidance work was separate from diagnostic foundation models, and sequential baselines typically fused image and action features only in the prediction head; this work combines diagnostic foundation-model image representations with sequence-based vision-action interaction, learning structure progressively during feature extraction. On 1.31M image-action pairs, EchoCLIP+Ours reaches average MAE of 5.40 mm translation and 6.74 degrees rotation with 2.61M trainable parameters, better than fully fine-tuned EchoCLIP (6.56 mm, 7.66 degrees) and other baselines.
It is markedly more parameter-efficient than full fine-tuning and other PEFT methods (LoRA, Prefix Tuning) while keeping real-time inference. Trainable parameters drop about 33-fold versus full fine-tuning baselines (e.g., EchoCLIP 88.41M versus 2.61M); ablation shows that at adapter dimension r=8 only 0.2M parameters (0.23% of full fine-tuning) already reduce MAE by 11.9%/8.2%. Table 1 lists trainable parameters and MAE for each baseline; on an RTX 3090 a single sequence takes about 10.0-11.1 ms without VA-Adapter and 10.5-11.8 ms with it.
The vision-action interaction module is the key source of the gain, not merely extra parameters. Compared with a vanilla adapter of similar scale, it reduces translation/rotation MAE by 12.6%/8.0% with only 0.9M extra parameters, and outperforms other methods after the first epoch. Ablation over seven training baselines for EchoCLIP (Fig. 4) and adapter-dimension ablation (Fig. 5).
The model is robust to cardiac-cycle variation and to non-standard, low-quality planes, inferring correct actions from vision-action relationships. Outputs are consistent across frames from the same probe position but different cardiac phases (systole/diastole); for non-standard planes with weak visual cues the sequence model still infers correct actions, which single-frame models cannot. Visualization in Fig. 6: applying the predicted action to obtain a pose, then nearest-neighbor retrieval in the scan sequence, yields a retrieved plane matching the target.
Perspective
The result targets probe guidance for 10 standard planes in adult transthoracic echocardiography, with data from 178 adult subjects and 356 expert scanning trajectories collected by two senior sonographers using a GE machine and M5S probe under university medical ethics committee approval and supervision; 284 scans are used for training and 72 for validation, with different subjects in each set. The method applies to ultrasound foundation models that already have an image encoder (EchoCLIP, USFM, BiomedCLIP), attached as an adapter while the backbone stays frozen, so it suits settings that want low-cost reuse of existing foundation models and have pose-annotated scan sequences, for example real-time assistance to junior sonographers or as the decision-making core of autonomous robotic ultrasound systems.
Validation uses subjects different from training, but at 72 scans, generalization across centers, devices and populations still needs more data; the dataset covers only 10 standard planes in adult transthoracic scanning, so applicability to fetal, pediatric or transesophageal settings is unclear; the visualization demonstrates guidance by comparing a nearest-neighbor retrieved plane with the target rather than a closed-loop robotic clinical endpoint; and beyond inference latency the text does not report failure-mode statistics, so behavior under severe anatomical variation or very poor image quality remains an open question.
