MEF-TBN: A Three-Branch Network for Multi-Exposure Image Fusion with Feature Extraction and Color Enhancement
Synopsis
The work proposes MEF-TBN, a three-branch network operating in YCbCr space that uses a context aggregation attention network (CAAN) branch for multi-scale local features, a transformer branch with multi-head self-attention for global long-range dependencies, and a color enhancement branch that learns the mapping between luminance and chrominance; the two feature branches produce low-resolution weight maps that are refined by a merger module with three merger blocks and upsampled to original resolution via guided filtering for joint upsampling (GFU), and across three MEF datasets against 9 representative methods it achieves the best results on image-feature metrics such as AG, EI and SF and on all human-perception metrics such as QCB, VIF and NIQE, ranks second on MEF-SSIM, and improves ove
Interpretation
It assigns local multi-scale feature extraction to a CAAN branch and global long-range dependency modeling to a transformer branch, then integrates the two low-resolution weight maps through a merger module (three merger blocks) using element-wise addition plus convolution-based refinement, and finally produces original-resolution weight maps via guided filtering for joint upsampling (GFU). Compared with earlier hybrid CNN-Transformer methods that rely on simple concatenation or complex attention-based fusion, this design preserves the original representations of both branches and refines them with convolutions; the authors report that removing any branch markedly lowers average MEF-SSIM, QCB and VIF, and that replacing the merger with cross-attention-style fusion drops these three metrics to 0.9711, 0.5138 and 0.9026. Trained on the SICE set (589 static sequences, each with at least 3 exposures), evaluated quantitatively and subjectively on three test datasets including MEFB, with ablation tables covering the three branches, the merger module, the upsampling scheme, resolution, and network depth and width.
It introduces a color enhancement branch that takes the most over-exposed and most under-exposed images together with the fused Y channel as input and uses a symmetric encoder-decoder fully convolutional network to learn the luminance-chrominance mapping, replacing simple weighted fusion of Cb and Cr. Addressing the common practice of designing a fusion strategy only for the luminance component while applying simple weighted summation to chrominance, which causes color distortion in over-exposed or under-exposed regions, this branch directly models the color-luminance relationship; ablation shows the average Colorfulness Index drops markedly when the branch is removed. Ablation uses MEF-SSIM, QCB, VIF and the Colorfulness Index with supporting subjective comparisons; during training the most over-exposed, under-exposed and middle-exposed images are selected from each sequence for this branch.
It reports a unified comparison against 9 representative MEF methods, including 10 metrics, a user preference study and an efficiency comparison. The authors report best results on EN, AG, EI, SF, SD, QCB, VIF and NIQE and second place on QY and MEF-SSIM; in blind pairwise preference comparisons by 8 volunteers (4 researchers and 4 non-professional participants) the method reaches the highest preference rate of 86.1% versus 73.1% for UltraFusion; the model uses 23.5M parameters, 26.7G FLOPs and 94MB, below SwinFusion's 34.2M, 41.5G and 136.8MB and UltraFusion's 217M, 126.3G and 868MB. Quantitative results are given as mean ± standard deviation with 95% confidence intervals, with Wilcoxon signed-rank tests (p<0.05) for the main metrics; some baselines were retrained on the same SICE data while others were evaluated with official pretrained models.
Perspective
The results target multi-exposure fusion of static scenes with well-aligned inputs; the method processes images in YCbCr space and supports arbitrary resolution and more than two inputs. The authors report stable behavior when trained at 512s and tested from 512s to 2048s, making it applicable to high-resolution static sequences and to photography or image enhancement pipelines that need natural color and human-perceptual quality.
The authors note that the current framework lacks motion detection or image alignment, so dynamic scenes produce ghost artifacts and future work should add alignment or deghosting modules. In addition, parts of the equations, table values and figure captions are truncated in the loaded text, for example the exact training resolution, the guided filter radius and some ablation values cannot be fully verified, so readers who need to reproduce the work should confirm these details against the original figures and tables.
