Refining photogrammetric DSMs with a pretrained diffusion model and multimodal conditioning cut Dense Urban RMSE from 6.00 m to 3.45 m in French cities
Synopsis
The study adapts pretrained Stable Diffusion 3 into an image-only generative backbone with a pruned text stream and patch-wise normalization, conditioning on both photogrammetric DSMs and Pléiades-HR imagery via two ControlNets to refine vertically co-registered DSMs, reducing Dense Urban RMSE from 6.00 m to 3.45 m across eight in-context French cities and from 4.16 m to 2.77 m in the geographically held-out city of Bordeaux.
Interpretation
Patch-wise normalization lets a pretrained diffusion model adapt stably to 3D elevation maps whose mean elevation and local relief vary widely, and restores predictions to their original physical units after sampling. Prior diffusion work on terrain focused mainly on void filling or text-driven generation, and direct flow matching on raw elevation is unstable because local patch variance is small relative to the dataset-wide elevation range; this work reparameterizes noise and target with scale and offset estimated from the conditioning DSM, normalizing the loss to roughly the scale used for natural-image training. The paper reports a minimum raw standard deviation of 0.00473 m over the scanned LiDAR support, uses no clipping or zero-variance fallback, and reports no instability attributable to patch-wise scaling in the reported training runs; the statistics are given in Appendix C.
Pruning the text stream turns SD3 into an image-only generator that keeps pretrained image weights while cutting the transformer from about 2B to 1B parameters and raising inference throughput. Relative to the original SD3 multimodal diffusion transformer, the work converts joint text-image attention into image self-attention and removes text-specific projections, feed-forward networks, LayerNorm and modulation layers, yielding a lighter backbone for a map-generation task with no textual annotations. Appendix B, Table A3 reports a synthetic 50-step denoising plus VAE-decoding benchmark: the pruned transformer has 1.033B parameters and each ControlNet 0.547B, with throughput about 1.74 samples/s at batch size 32 (about 0.82 with two ControlNets), above the unpruned configuration.
Conditioning on both the photogrammetric DSM and Pléiades RGB imagery lowers elevation error further than DSM-only conditioning, with the clearest gains in Dense Urban and Roads. The paper compares DSM-only and DSM + RGB under the same paired protocol and reports per-land-cover MAE and RMSE for both in-context cities and the held-out city of Bordeaux, showing RGB adds gains in most land-cover groups. On 6385 in-context geographic patches, Dense Urban RMSE falls from 6.00 m for the calibrated input to 3.85 m with DSM-only and 3.45 m with DSM + RGB; on 3209 Bordeaux patches it falls from 4.16 m to 3.10 m and 2.77 m, with metrics averaged over 20 inference seeds on common finite LiDAR/DSM support.
The refinement model has a secondary ability to recover original DSM voids, with DSM + RGB showing lower void RMSE than DSM-only in all five land-cover groups in both test sets. Void filling has largely been handled by dedicated DEM void-filling methods; here it is evaluated separately as a by-product of systematic correction of a dense stereo DSM, and the calibrated input has no numerical baseline at those locations. Appendix D.2, Table A7 reports void-pixel RMSE by land-cover group, including only pixels with valid LiDAR elevation but no input DSM elevation and excluding missing pixels outside the source DSM image.
Perspective
The work targets inputs satisfying two assumptions: the DSM and LiDAR are vertically co-registered (A1), and they represent the same underlying physical surface (A2); the estimand is correction of local errors, not general DSM-to-LiDAR reconstruction from an arbitrary, uncalibrated or surface-inconsistent input. The applicable setting is 0.5 m resolution Pléiades-HR and CARS stereo DSM pairs over French cities, with Bordeaux the single geographically held-out city within the same acquisition and processing setting. For a reader, two extension lines stand out: patch-wise normalization as a general means of adapting models to physical-valued rasters, and the pruned diffusion backbone as a pretrained representation for downstream geospatial tasks such as segmentation and object detection; the authors also propose assessing whether refining pre- and post-disaster stereo DSMs improves elevation-change maps, which would first require checking that refinement preserves genuine changes rather than suppressing or introducing them.
The paper itself flags several open questions: the RMSE filter is only an indirect compatibility heuristic rather than a verified surface-change mask, and it changes the land-cover composition of the evaluated data (Dense Urban pixels are more common by 1.81 among the exclusions); acquisition-date differences and residual misalignment complicate interpretation of DSM-LiDAR differences; the effect of interpolating small voids in the LiDAR reference around roofs and canopies remains unquantified; high-rise buildings, bridges and steep terrain are not evaluated separately; adjacent training and test patches can touch without a spatial buffer, so city-level bootstrap intervals do not remove spatial dependence between splits; and the 10 m Theia labels provide only coarse strata on the 0.5 m grid. In addition, the experiments establish gains over the calibrated DSM and the DSM-only configuration rather than superiority over other enhancement methods, and task-matched conventional and supervised baselines, alternative fusion and training strategies, and independently repeated training runs remain to be added. This summary is based on the full paper text, but the source imagery, elevation rasters and model weights were not obtained, as they cannot be redistributed under the applicable licenses.
