On a 640-KB-SRAM STM32H563ZI MCU, an INT8-quantized super-resolution network simultaneously reconstructs high-resolution depth and intensity from a VL53L9CX SPAD sensor
Synopsis
This work presents a two-branch super-resolution deep-learning framework that simultaneously upscales low-resolution depth and intensity images from a consumer-grade VL53L9CX SPAD time-of-flight sensor, using a two-stage PixelShuffle reconstruction head and simplified sensor firmware to control activation memory, then exports the model to ONNX, quantizes it to INT8, generates C code, and deploys it on a bare-metal STM32H563ZI MCU with only 640 KB SRAM and 2 MB Flash for fully on-device super-resolution inference.
Fig. 1: (a) DL architecture. Modality-specific branches process depth and intensity, with intensity-gradient guidance provided to the depth branch. (b) Simplified PixelShuffle module. The interchangeable backbone is (c) SAFM, (d) RLFN, or (e) SPAN. Each branch upsamples through two compact PixelShuffle stages and uses an interpolated residual connection for reconstruction.
arXivInterpretation
A two-branch super-resolution network is proposed and validated for simultaneous depth and intensity reconstruction, where the depth branch combines Charbonnier, Sobel-gradient, total-variation, and structural loss terms, while the intensity branch omits total variation because reflectance texture is signal rather than artifact for intensity. Unlike prior methods that fuse histograms with high-resolution RGB and run on GPUs or large-memory embedded platforms, this framework takes only single-sensor low-resolution depth and intensity as input and requires no cross-camera registration. On eight Middlebury synthetic scenes, all three learned models reduce depth RMSE from 0.051 m for bicubic interpolation to 0.042 m, an 18% improvement; intensity PSNR rises from 26.88 dB to 27.71 dB (SAFM), and GMSD drops from 0.0017 to 0.0014.
A memory-efficient two-stage PixelShuffle reconstruction head is designed by splitting the single-stage head into two stages and reducing intermediate channels, bringing peak activation memory within SRAM capacity and cutting parameter count from 21,037 to 3,780 INT8 numbers. Whereas super-resolution memory bottlenecks are usually attributed to weight storage, this work identifies peak SRAM occupancy as dominated by large activation tensors after the backbone and redesigns the upscaling modules accordingly. The three quantized backbones use 522 to 533 KB of RAM on the MCU (about 81% to 83%), 141 to 171 KB of Flash (about 7% to 8%), and achieve inference latencies of 0.44 to 0.75 ms.
An automated deployment pipeline is established from PyTorch training through ONNX export, INT8 post-training static quantization, ST Edge AI Core compilation to C, and integration with simplified sensor firmware, enabling on-device super-resolution on a bare-metal STM32H563ZI MCU. According to the authors, this is the first non-fusion super-resolution deep-learning model running entirely on a resource-constrained MCU that reconstructs depth and intensity maps directly from measurements of a consumer-grade low-resolution SPAD sensor. After INT8 quantization, intensity PSNR decreases only from 27.71 to 27.68 dB and SSIM from 0.8226 to 0.8213; depth RMSE increases from 0.0423 m to 0.0470 m and AbsRel from 0.0093 to 0.0158.
In real measurements, the super-resolution output clearly resolves the human outline, the gap between arm and torso, the raised hand, and facial and clothing texture that appear as jagged silhouettes and flat bright patches at low resolution, and the point cloud separates the person from the wall behind. The real-data validation complements the synthetic experiments, showing that the framework still produces denser depth, intensity, and point-cloud outputs under the actual noise and scene conditions of a consumer-grade SPAD sensor. Seven real captures (an empty scene and a person walking from close to far and back) show that the high-resolution depth gradient is no longer striped and that facial features and clothing texture emerge from flat bright patches in the intensity channel.
Perspective
The result targets joint depth and intensity super-resolution on a bare-metal STM32H563ZI MCU (single-core Arm Cortex-M33, 250 MHz, 640 KB SRAM, 2 MB Flash) with a VL53L9CX consumer-grade SPAD time-of-flight sensor as input, at a total system cost of about USD 120. It enables low-cost robotics, autonomous platforms, and 3D perception devices to obtain denser depth, intensity, and point-cloud outputs without a GPU, embedded Linux processor, or dedicated neural processing unit. The authors note that further gains could come from more efficient backbone architectures, compilation and execution pipeline optimization (such as target-specific DSP instruction-set kernels and activation-buffer management), and restricting high-resolution reconstruction to spatial regions that genuinely need detail (region-of-interest super-resolution).
The synthetic evaluation uses only eight Middlebury scenes and the real evaluation has seven captures, so the sample scope is limited and behavior across broader scenes and sensor batches remains unclear. Depth threshold metrics are nearly saturated, and the authors note these metrics provide limited discrimination over the depth range considered. INT8 quantization introduces measurable degradation in absolute depth accuracy (RMSE from 0.0423 m to 0.0470 m, AbsRel from 0.0093 to 0.0158) while intensity reconstruction is nearly unaffected, an asymmetry worth watching in accuracy-sensitive applications. The loss-weight ablation shows most weights change performance little over an order of magnitude, and the authors suggest further joint tuning of a few weights. Region-of-interest super-resolution is proposed to reduce latency, but the trade-off between ROI-selection accuracy and computational savings has not been evaluated. In addition, several numeric values in the loaded text (such as SRAM capacity, loss weights, latencies, and memory byte counts) are missing as placeholders, so exact reproduction should consult the original figures and tables.
