FloVMos fine-tunes optical-flow networks on synthetic data to deliver real-time medical video mosaicking across seven imaging modalities, outperforming three baselines
Related research and updatesSynopsis
The authors present FloVMos, an optical-flow-based deep learning video mosaicking framework that fine-tunes an optical flow network on synthetic training data with ground-truth deformation fields, generated automatically from existing large-FOV mosaics or raw videos to adapt to different imaging modalities; across seven modalities (reflection confocal microscopy, open-top light-sheet microscopy, fetoscopy, laparoscopy, dermoscopy, sparse spectral microscopy, and endoscopy), FloVMos outperformed the Parallax, APAP, and AVM baselines in accuracy, robustness, and speed, preserving feature-point distances with an average error of about 0.5% of the field of view and processing 1 MP images at roughly ten frames per second.
Figure 1: An overview of the FloVMos videomosaicking pipeline highlighting its four key steps: (i) data acquisition, (ii) image registration using a fine-tuned optical flow estimator (a deep neural network), (iii) real-time image warping, and (iv) image alignment and blending.
arXivInterpretation
FloVMos uses an optical flow deep network for pixel-level registration and registers each frame directly to a region of interest on the growing mosaic rather than recursively frame-to-frame, limiting error accumulation. Prior methods handling non-rigid deformation were specialized to specific modalities and acquisition pathways and computationally costly, leaving no general, fast, easily adaptable solution. The paper describes a four-step pipeline (acquisition, optical flow registration, warping, graph-cut blending) and reports about ten frames per second on 1 MP images in its research code implementation.
The authors introduce two annotation-free routes for synthetic training data: simulating a limited-FOV device path over existing large-FOV mosaics, or, for video-only modalities, tiling a single frame with flips and arranging four frames that are at least 15% of the video apart with under 25% overlap. This removes the need for labor-intensive pixel-level optical flow annotation and enables rapid per-modality fine-tuning of the optical flow model. The paper provides per-modality parameters for translation, rotation, shear, and non-rigid level, set from prior knowledge, such as non-rigid level 3 for dermoscopy and 10 for RCM.
Across seven modalities, FloVMos preserves feature-point distances after registration with an average error of about 0.5% of the field of view; on simulated RCM data its SSIM and LPIPS mean values and distributions outperform Parallax, APAP, and AVM, with the difference versus the best competitor, Parallax, statistically significant (p-value 0.0001). Comparison methods were limited under non-rigid deformation and long videos, whereas FloVMos improved both accuracy and stability. Comparisons used simulated RCM data with available ground truth, two standard metrics (SSIM and LPIPS), and full metric distributions.
FloVMos is modular and robust to the choice of optical flow network: after fine-tuning on the same data, FlowNet2, RAFT, FlowFormer, and GMFlow all showed large MSE and F1 reductions, e.g., FlowNet2 MSE dropping from 0.940 to 0.046. This indicates the framework's gains come mainly from synthetic-data fine-tuning and the overall pipeline rather than from binding to one specific optical flow network. All four networks were fine-tuned on the same 11,000 videos with 40,000 frames and evaluated on a distinct set of 172 videos with 6,000 frames.
Perspective
The work targets close-proximity modalities where the imaging device is near or in contact with tissue, and applies to research and clinical imaging workflows seeking low-cost field-of-view extension; for modalities with existing large-FOV mosaics or raw videos, users can automatically generate training data and fine-tune the optical flow network to obtain real-time mosaicking on a GPU. The paper notes that when only raw videos are available, attainable dataset sizes are substantially smaller than typical deep learning benchmarks, and increasing synthetic dataset size and diversity helps mitigate uncertainty from the absence of direct temporal reference; when real mosaics from the same modality exist, larger synthetic datasets can be generated. The authors also report robust estimation with at least 50% overlap between consecutive frames, and that selective downsampling can accelerate processing when high-frame-rate videos have greater overlap.
The paper lists reliance on simulated data as a major limitation of the current implementation and notes the study concentrated on close-proximity modalities; devices farther from tissue may introduce distortions not covered by the current simulation model, so future work should test on real-world clinical data and fine-tune as needed. Real-time processing requires a GPU, which most clinical imaging systems do not meet; the authors suggest CPU-efficient models or compression techniques such as pruning, quantization, and knowledge distillation, but the trade-offs remain uncertain. Intraoperative use currently rests on early feasibility studies, and effects on surgical outcomes and workflow efficiency still need evaluation. In addition, parameters such as the non-rigid deformation level depend on domain prior knowledge, and when priors are limited, a small synthetic set with varying deformation levels can be qualitatively compared against real modality data. Readers should also note that this is a full-text parse in which specific value distributions in figures and some formula symbols are not fully rendered in the text, so those details require consulting the original figures.
