Skip to main content
Back to timeline
arXivSource publication:

Flow matching for sparse-view CT reconstruction: reusing velocity fields cuts network evaluations to 7 while matching diffusion-model quality

Synopsis

The authors propose FMCT, the first flow-matching-based CT reconstruction framework, and an efficient variant EFMCT that reuses previously predicted velocity fields over consecutive steps to cut neural network function evaluations (NFEs), motivated by the deterministic ODE sampling of flow matching being naturally compatible with repeated data consistency corrections and by the strong correlation of predicted velocities across adjacent steps; they provide a theoretical analysis showing the reuse error is of the same order as Euler discretization and bounded for a bounded number of consecutive reuse steps, and on the AAPM and Medical Segmentation Decathlon CT datasets with 20-view and 40-view parallel-beam sparse sampling, FMCT/EFMCT achieve PSNR/SSIM competitive with diffusion-based method

Source-provided article image: Efficient Flow Matching for Sparse-View CT Reconstruction

Interpretation

They propose FMCT, the first flow-matching-based framework for CT reconstruction, combining deterministic ODE sampling with physics-based data consistency for the underdetermined sparse-view CT inverse problem. Sparse-view CT generative reconstruction has been dominated by diffusion models, whose forward diffusion and reverse denoising are modeled as stochastic differential equations; that stochasticity interferes with the repeated data consistency corrections, producing a push-and-pull effect. Using deterministic ODE transport yields smooth trajectories that interact more coherently with repeated corrections. The paper derives the method (Euler update, linear extrapolation to estimate the terminal sample, conjugate gradient for data consistency) and compares against FBP, ADMM-TV, DPS, MCG, PGDM, and DDS on the AAPM and Decathlon datasets under 20-view and 40-view parallel-beam settings.

They introduce a velocity reuse strategy: after computing the velocity at time t, the same velocity is reused for up to M consecutive steps, with an adaptive data-consistency residual criterion deciding whether reuse continues, directly reducing NFEs. The strategy rests on the observed strong correlation of predicted velocities across adjacent steps (cosine similarity in Fig. 1), treating per-step network re-evaluation as redundant; unlike approaches that only reduce the number of reverse steps via deterministic samplers, it reduces network evaluations per step rather than the step count itself. Ablations on 21 randomly selected test CT images with total sampling iterations fixed at 50 show that enabling reuse too early first degrades performance while delaying it gradually improves PSNR/SSIM, and that increasing the reuse limit first helps then degrades, revealing a quality-efficiency trade-off. All EFMCT results enable reuse after the first iteration with up to 10 consecutive reuse steps and relaxation factor η = 1.05.

They provide a theoretical analysis: the single-step reuse deviation from the standard Euler update is O(Δt²), the same order as Euler discretization, and with bounded reuse length M and a data-consistency correction that is non-expansive in its image argument, the accumulated deviation after one reuse block is O(M²Δt). This gives a provable basis that velocity reuse does not break convergence, rather than only an empirical speed-up trick, linking the efficiency gain to error control. Proposition 1 is proved under local Lipschitz continuity of the velocity field in x and t and bounded update increments, with the full derivation in the appendix.

On two public CT datasets, FMCT/EFMCT are often among the highest PSNR and SSIM across most datasets and view configurations while requiring substantially fewer NFEs and lower computation time than diffusion-based baselines. Relative to the most efficient diffusion baseline, FMCT uses 25 NFEs and 1.92 s on AAPM 40 views (about 50% reduction) and EFMCT uses 7 NFEs and 0.83 s (about 75% and 78% reductions); on Decathlon EFMCT uses 11 NFEs and 1.72 s (about 89% and 78% reductions). Compared with FMCT, EFMCT shows a slight reduction in PSNR and data fidelity but comparable or improved SSIM. Table 1 reports means and standard deviations for PSNR, SSIM, and data fidelity; Table 2 reports LPIPS; Fig. 3 gives visual comparisons across datasets and views and notes that ADMM-TV, despite high PSNR/SSIM, is overly smooth and lacks fine details.

Perspective

The work targets sparse-view CT reconstruction. Experiments use the AAPM 2016 Low Dose CT Grand Challenge dataset (nine patients for training, patient L506 for testing) and the Medical Segmentation Decathlon Task_06 Lung subset (ten patients for training, patient 017 for testing), with 20-view and 40-view parallel-beam projections simulated using the ASTRA Toolbox; the authors state that parallel-beam geometry was chosen deliberately to provide a controlled setting that minimizes confounding from geometric complexity and isolates the reconstruction methodology. Methodologically, data consistency is implemented with conjugate gradient, though the framework is stated to be compatible with other data consistency schemes; the authors also note FMCT/EFMCT do not rely on geometry-specific assumptions and can be extended to other acquisition geometries by modifying the forward operator. The results therefore apply to research and engineering settings using simulated sparse-view parallel-beam data and concerned with the trade-off between inference efficiency and reconstruction quality; the codebase is open-sourced.

A careful reader would still watch whether the efficiency advantage of velocity reuse holds on real scan data, other acquisition geometries (e.g., cone-beam, limited-angle), and under noise; whether the slight drop in PSNR and data fidelity of EFMCT relative to FMCT is acceptable in clinical reading; whether the optimal reuse limit M and enabling time depend on dataset and view count; and how well the Lipschitz and bounded-increment assumptions behind the theoretical analysis hold for actually trained models. In addition, this is a full-text parse, and the specific numerical curves of Fig. 1 and Fig. 4 are not given point by point in the text, so exact similarity and ablation values still require consulting the original figures.

Sources