DenseFuse: A Fusion Approach to Infrared and Visible Images

Hui LiXiaojun Wu

article2018IEEE Transactions on Image Processing1,935 citationsHighly Cited Paper (IEEE TIP)

Proposes an autoencoder framework incorporating dense blocks and custom fusion strategies to effectively capture and merge complementary features from infrared and visible images for superior visual reconstruction.

Listen

Combining visible and infrared imaging is essential for operational environments such as nighttime video surveillance and defense applications. Visible imagery captures rich environmental textures and fine background details, whereas infrared imagery detects thermal signatures in low-visibility conditions. Conventional fusion techniques frequently blur fine details, introduce artificial visual noise, or discard valuable intermediate features discovered by deep learning models. The article sets out to design, implement, and validate DenseFuse, a deep learning architecture that preserves multi-layer visual features to generate clearer, higher-fidelity fused images.

The researchers developed a modular framework consisting of an encoding network with densely connected convolutional layers, a fusion layer, and a four-layer decoding network. To train the encoder and decoder to extract and reconstruct salient features effectively, the system utilized roughly 80,000 visible images from the MS-COCO dataset with a loss function balancing structural similarity and raw pixel accuracy. The researchers evaluated the trained model across 20 registered image pairs, testing both a simple addition strategy and a sophisticated l1-norm saliency-based fusion strategy against six established traditional and deep learning algorithms across seven objective image-quality metrics.

The evaluation produced several significant findings. First, the DenseFuse architecture achieved top-tier performance across all tested objective metrics, achieving the highest average scores in entropy, structural similarity preservation, and feature mutual information. Second, visual inspections demonstrated that the proposed method retained clear foreground targets and sharp background textures while minimizing the artificial noise and excessive darkening seen in existing approaches. Third, incorporating dense blocks successfully allowed intermediate layer features to flow through the network, preventing the degradation and feature loss common in conventional neural networks. Finally, adjusting structural similarity loss weights accelerated network convergence without sacrificing final image reconstruction quality.

These results demonstrate that dense feature-reuse networks can substantially improve the visual and quantitative fidelity of multi-modal imagery. For mission-critical operations such as automated target tracking and perimeter defense, cleaner image fusion reduces human operator fatigue and lowers false-alarm risks. The decoupled training design also provides high operational flexibility: the base network requires training only once, allowing different fusion strategies to be deployed for grayscale or color imaging tasks without retraining the core model.

Organizations developing computer vision systems should consider adopting dense feature-reuse architectures for multi-source image processing. System engineers can choose the straightforward addition strategy for low-complexity deployments or the l1-norm strategy when edge preservation and structural detail are paramount. Further research and piloting should explore adapting this architecture to specialized operational domains, including medical diagnostics, multi-exposure photography, and multi-focus imaging.

Decision-makers should note that the evaluation was conducted on a relatively small benchmark of 20 image pairs and assumes that all incoming image pairs are accurately pre-aligned and calibrated. The primary training was also performed using visible natural scenes due to limited public infrared data. Nonetheless, the high consistency across subjective reviews and multiple quantitative benchmarks provides strong confidence in the architecture's core capabilities.

Cover for DenseFuse: A Fusion Approach to Infrared and Visible Images

Abstract

In this paper, we present a novel deep learning architecture for infrared and visible images fusion problem. In contrast to conventional convolutional networks, our encoding network is combined by convolutional layers, fusion layer and dense block in which the output of each layer is connected to every other layer. We attempt to use this architecture to get more useful features from source images in encoding process. And two fusion layers(fusion strategies) are designed to fuse these features. Finally, the fused image is reconstructed by decoder. Compared with existing fusion methods, the proposed fusion method achieves state-of-the-art performance in objective and subjective assessment. Code and pre-trained models are available at this https URL

Table of Contents

  • I Introduction
  • II Related Works
  • III Proposed Fusion Method
  • III-A Training
  • III-B Fusion Layer(strategy)
  • III-B1 Addition Strategy
  • III-B2 l1l_{1}-norm Strategy
  • IV Experimental results and analysis
  • IV-A Training phase analysis
  • IV-B Experimental Settings
  • IV-C Fusion methods Evaluation
  • IV-D Additional results for RGB images and infrared images
  • V Conclusion
  • References

Knowls

  1. Knowl 1 — DenseFuse Autoencoder Architecture for Image Fusion

    model/method

    DenseFuse is a deep learning framework designed for infrared and visible image fusion comprising three main functional blocks: an encoder, a fusion layer, and a decoder.

    The encoder extracts deep multi-layer salient features and consists of an initial convolutional layer (C1) and a Dense Block with three subsequent convolutional layers (DC1, DC2, DC3):

    • Layer C1 contains sixteen 3×33 \times 3 filters with stride 1, ReLU activation, and reflection padding, transforming a 1-channel grayscale input into a 16-channel feature map.
    • The Dense Block connects the output of each convolutional layer to all subsequent layers via channel concatenation. Layer DC1 takes the 16-channel output of C1 and produces 16 channels with 3×33 \times 3 filters (stride 1, ReLU). Layer DC2 takes the concatenated 32-channel output of C1 and DC1 and produces 16 channels (3×33 \times 3, stride 1, ReLU). Layer DC3 takes the concatenated 48-channel output of C1, DC1, and DC2 and produces 16 channels (3×33 \times 3, stride 1, ReLU).
    • The total output of the encoder concatenates the outputs of C1, DC1, DC2, and DC3, yielding a 64-channel feature map (M=64M = 64).

    The fusion layer merges the 64-channel feature maps from multiple input images into a single 64-channel feature map.

    The decoder reconstructs the fused image from the 64-channel merged representation using four successive convolutional layers with 3×33 \times 3 filters, stride 1, and reflection padding:

    • Layer C2: 64 input channels to 64 output channels with ReLU.
    • Layer C3: 64 input channels to 32 output channels with ReLU.
    • Layer C4: 32 input channels to 16 output channels with ReLU.
    • Layer C5: 16 input channels to 1 output channel without activation.

    Because all convolutional layers have a stride of 1 and use padding without pooling or downsampling, DenseFuse accepts input source images of arbitrary spatial dimensions.

  2. Knowl 2 — DenseFuse Training Objective and Composite Loss Function

    equation

    In the training phase of DenseFuse, the encoder and decoder are trained end-to-end as an unsupervised autoencoder to reconstruct input visible images, setting aside the fusion layer. The network weights are optimized by minimizing a composite loss function LL:

    L=λLssim+LpL = \lambda L_{ssim} + L_p

    where LpL_p is the pixel-level Euclidean reconstruction loss:

    Lp=∥O−I∥2L_p = \|O - I\|_2

    and LssimL_{ssim} is the structural dissimilarity loss:

    Lssim=1−SSIM(O,I)L_{ssim} = 1 - SSIM(O, I)

    Here, II denotes the input grayscale image, OO denotes the reconstructed output image produced by passing II through the encoder and decoder, SSIM(O,I)SSIM(O, I) is the structural similarity index measure evaluating luminance, contrast, and structural preservation, and λ>0\lambda > 0 is a scalar hyperparameter weighting the SSIM loss against the Euclidean pixel loss to compensate for the numerical difference of several orders of magnitude between the two terms.

  3. Knowl 3 — Activity-Level Fusion Strategy via L1-Norm and Softmax Normalization

    algorithm

    The ℓ1\ell_1-norm and softmax fusion strategy computes adaptive spatial weight maps based on local deep feature activity levels to fuse feature tensors extracted by the DenseFuse encoder.

    Input: Registered source images I1,I2,…,IkI_1, I_2, \dots, I_k, trained DenseFuse encoder E\mathcal{E}, trained DenseFuse decoder D\mathcal{D}, spatial neighborhood radius r=1r = 1.
    Output: Fused image IfI_f.
    for each source image index i∈{1,…,k}i \in \{1, \dots, k\} do
        ϕi←E(Ii)\phi_i \leftarrow \mathcal{E}(I_i) where ϕi∈RH×W×M\phi_i \in \mathbb{R}^{H \times W \times M} with M=64M = 64
        for each spatial coordinate (x,y)(x, y) do
            Ci(x,y)←∑m=1M∣ϕim(x,y)∣C_i(x, y) \leftarrow \sum_{m=1}^M |\phi_i^m(x, y)|
        end for
    end for
    for each source image index i∈{1,…,k}i \in \{1, \dots, k\} do
        for each spatial coordinate (x,y)(x, y) do
            C^i(x,y)←1(2r+1)2∑a=−rr∑b=−rrCi(x+a,y+b)\hat{C}_i(x, y) \leftarrow \frac{1}{(2r + 1)^2} \sum_{a=-r}^r \sum_{b=-r}^r C_i(x + a, y + b)
        end for
    end for
    for each spatial coordinate (x,y)(x, y) do
        S(x,y)←∑n=1kC^n(x,y)S(x, y) \leftarrow \sum_{n=1}^k \hat{C}_n(x, y)
        for each source image index i∈{1,…,k}i \in \{1, \dots, k\} do
            wi(x,y)←C^i(x,y)S(x,y)w_i(x, y) \leftarrow \frac{\hat{C}_i(x, y)}{S(x, y)}
        end for
        for each channel m∈{1,…,M}m \in \{1, \dots, M\} do
            fm(x,y)←∑i=1kwi(x,y)⋅ϕim(x,y)f^m(x, y) \leftarrow \sum_{i=1}^k w_i(x, y) \cdot \phi_i^m(x, y)
        end for
    end for
    If←D(f)I_f \leftarrow \mathcal{D}(f)
    return IfI_f

    In this algorithm, Ci(x,y)=∥ϕi1:M(x,y)∥1C_i(x, y) = \|\phi_i^{1:M}(x, y)\|_1 serves as an initial spatial activity indicator across all M=64M=64 feature channels at coordinate (x,y)(x, y). The local average pooling with radius r=1r=1 over a (2r+1)×(2r+1)=3×3(2r+1) \times (2r+1) = 3 \times 3 window smooths the activity map into C^i(x,y)\hat{C}_i(x, y). Spatial weights wi(x,y)w_i(x,y) are obtained via softmax-like normalization across the kk input sources, and the fused feature map ff is reconstructed into image IfI_f by the decoder D\mathcal{D}.

  4. Knowl 4 — Addition-Based Feature Fusion Strategy

    model/method

    In the addition fusion strategy for DenseFuse, once the encoder and decoder network weights are fixed from autoencoder pre-training, deep feature representations are extracted for each of k≥2k \ge 2 registered source images IiI_i (i∈{1,…,k}i \in \{1, \dots, k\}) using the encoder:

    ϕi=Encoder(Ii)∈RH×W×M\phi_i = \text{Encoder}(I_i) \in \mathbb{R}^{H \times W \times M}

    where M=64M = 64 represents the total number of encoder feature channels. The fused feature map tensor f∈RH×W×Mf \in \mathbb{R}^{H \times W \times M} is formed by computing the element-wise sum across all source representations:

    fm(x,y)=∑i=1kϕim(x,y)f^m(x, y) = \sum_{i=1}^k \phi_i^m(x, y)

    where m∈{1,…,M}m \in \{1, \dots, M\} indexes the feature channel, and (x,y)(x, y) denotes spatial pixel coordinates. The fused feature map tensor ff is then passed as the direct input to the decoder network to reconstruct the final fused grayscale image.

  5. Knowl 5 — Multichannel Extension for Color Visible and Infrared Image Fusion

    model/method

    DenseFuse handles color (RGB) visible images and grayscale infrared images without modifying or retraining the underlying network architecture. The RGB visible image is separated into three individual grayscale color channels: Red (IvisRI_{vis}^R), Green (IvisGI_{vis}^G), and Blue (IvisBI_{vis}^B).

    Each color channel is paired separately with the single-channel infrared image IirI_{ir}, forming three distinct 2-image input pairs: (IvisR,Iir)(I_{vis}^R, I_{ir}), (IvisG,Iir)(I_{vis}^G, I_{ir}), and (IvisB,Iir)(I_{vis}^B, I_{ir}). Each channel pair is independently processed through the pre-trained DenseFuse encoder, fusion layer (utilizing either the addition strategy or the ℓ1\ell_1-norm strategy), and decoder, generating three fused grayscale output channels (IfR,IfG,IfBI_f^R, I_f^G, I_f^B).

    Finally, the three fused channel maps are concatenated along the color dimension to form the reconstructed color fused RGB image.

  6. Knowl 6 — Modified Average Structural Similarity Index for Image Fusion (SSIMa)

    definition

    The modified average structural similarity index (SSIMaSSIM_a) is an objective no-reference metric for evaluating how effectively a fused image preserves the structural information of two registered input source images. For a fused image FF and source images I1I_1 and I2I_2, SSIMaSSIM_a is defined as:

    SSIMa(F)=0.5×(SSIM(F,I1)+SSIM(F,I2))SSIM_a(F) = 0.5 \times \left( SSIM(F, I_1) + SSIM(F, I_2) \right)

    where SSIM(⋅,⋅)SSIM(\cdot, \cdot) denotes the standard structural similarity operation. Higher values of SSIMaSSIM_a correspond to superior structural retention and contrast fidelity from both source images.

  7. Knowl 7 — DenseFuse Autoencoder Training Setup

    experimental setup

    The DenseFuse encoder and decoder are trained in an unsupervised autoencoder configuration using visible images from the MS-COCO dataset. 80,000 images are resized to 256×256256 \times 256 pixels and converted to grayscale, with approximately 79,000 images assigned to the training set and 1,000 images reserved for validation.

    The network is optimized using a learning rate of 1×10−41 \times 10^{-4}, a batch size of 2, and 4 training epochs on an NVIDIA GTX 1080Ti GPU with a TensorFlow backend. The SSIM loss weight λ\lambda in the composite objective function L=λLssim+LpL = \lambda L_{ssim} + L_p is tested across four scale values: λ∈{1,10,100,1000}\lambda \in \{1, 10, 100, 1000\}.

  8. Knowl 8 — Quantitative Evaluation on 20 Infrared and Visible Image Pairs

    data/table

    The performance of DenseFuse with addition and ℓ1\ell_1-norm strategies across four SSIM weights (λ∈{100,101,102,103}\lambda \in \{10^0, 10^1, 10^2, 10^3\}) was evaluated against six existing fusion methods: Cross Bilateral Filter (CBF), Joint-Sparse Representation (JSR), Gradient Transfer and Total Variation Minimization (GTF), JSR with Saliency Detection (JSRSD), a CNN-based multi-focus method adapted to infrared/visible images (CNN), and DeepFuse. The evaluation was conducted across 20 infrared and visible image pairs using seven quality metrics: Entropy (EnEn), QAB/FQ^{AB/F} (QabfQabf), Sum of Correlations of Differences (SCDSCD), Wavelet Feature Mutual Information (FMIwFMI_w), Discrete Cosine Feature Mutual Information (FMIdctFMI_{dct}), Modified Structural Similarity (SSIMaSSIM_a), and Multi-Scale SSIM (MS_SSIMMS\_SSIM).

    Methods En Qabf SCD FMI_w FMI_dct SSIM_a MS_SSIM
    CBF 6.81494 0.44119 1.38963 0.32012 0.26619 0.60304 0.70879
    JSR 6.78576 0.32572 1.59136 0.18506 0.14184 0.53906 0.75523
    GTF 6.63597 0.40992 1.00488 0.41004 0.39384 0.70369 0.80844
    JSRSD 6.78441 0.32553 1.59124 0.18502 0.14201 0.53963 0.75517
    CNN 6.80593 0.29451 1.48060 0.53954 0.35746 0.71109 0.80772
    DeepFuse 6.68170 0.43989 1.84525 0.42438 0.41357 0.72949 0.93353
    DenseFuse (Addition, λ=100\lambda = 10^0) 6.66280 0.44114 1.84929 0.42713 0.41557 0.73159 0.93039
    DenseFuse (Addition, λ=101\lambda = 10^1) 6.65139 0.44039 1.84549 0.42707 0.41552 0.73246 0.92896
    DenseFuse (Addition, λ=102\lambda = 10^2) 6.65426 0.44190 1.84854 0.42731 0.41587 0.73186 0.92995
    DenseFuse (Addition, λ=103\lambda = 10^3) 6.64377 0.43831 1.84172 0.42699 0.41558 0.73259 0.92794
    DenseFuse (ℓ1\ell_1-norm, λ=100\lambda = 10^0) 6.83278 0.47560 1.71182 0.43191 0.38062 0.71880 0.85707
    DenseFuse (ℓ1\ell_1-norm, λ=101\lambda = 10^1) 6.81348 0.47680 1.71264 0.43224 0.38048 0.72052 0.85803
    DenseFuse (ℓ1\ell_1-norm, λ=102\lambda = 10^2) 6.83091 0.47684 1.71705 0.43129 0.38109 0.71901 0.85975
    DenseFuse (ℓ1\ell_1-norm, λ=103\lambda = 10^3) 6.84189 0.47595 1.72000 0.43147 0.38404 0.72106 0.86340

    The quantitative results indicate that DenseFuse with ℓ1\ell_1-norm fusion achieves the highest average Entropy (En=6.84189En = 6.84189 at λ=103\lambda = 10^3) and highest edge information retention (Qabf=0.47684Qabf = 0.47684 at λ=102\lambda = 10^2), while DenseFuse with the addition strategy achieves the highest correlation difference sum (SCD=1.84929SCD = 1.84929 at λ=100\lambda = 10^0), highest discrete cosine mutual information (FMIdct=0.41587FMI_{dct} = 0.41587 at λ=102\lambda = 10^2), and highest structural similarity (SSIMa=0.73259SSIM_a = 0.73259 at λ=103\lambda = 10^3).

  9. Knowl 9 — Impact of SSIM Loss Weight on Training Convergence and Validation Accuracy

    empirical result

    Empirical analysis of the loss weighting factor λ\lambda in L=λLssim+LpL = \lambda L_{ssim} + L_p during autoencoder training demonstrates that larger values of λ\lambda (e.g., λ=1000\lambda = 1000) result in significantly faster convergence during the initial 2,000 training iterations compared to lower values (λ=1\lambda = 1). On the 1,000-image validation set, higher λ\lambda values achieve superior SSIM scores and lower pixel reconstruction errors within the first 500 iterations.

    However, after extensive training exceeding 40,000 iterations, the autoencoder converges to comparable optimal weights and reconstruction performance across all tested values λ∈{1,10,100,1000}\lambda \in \{1, 10, 100, 1000\}. Thus, selecting a larger λ\lambda primarily accelerates early convergence and reduces overall training time.

Coverage note — No substantial contributed material was omitted.

References

  1. 1.Li S, Kang X, Fang L, et al. Pixel-level image fusion: A survey of the state of the art[J]. Information Fusion, 2017, 33: 100-112.
  2. 2.Ben Hamza A, He Y, Krim H, et al. A multiscale approach to pixel-level image fusion[J]. Integrated Computer-Aided Engineering, 2005, 12(2): 135-146.
  3. 3.Yang S, Wang M, Jiao L, et al. Image fusion based on a new contourlet packet[J]. Information Fusion, 2010, 11(2): 78-84.
  4. 4.Wang L, Li B, Tian L F. EGGDD: An explicit dependency model for multi-modal medical image fusion in shift-invariant shearlet transform domain[J]. Information Fusion, 2014, 19: 29-37.
  5. 5.Pang H, Zhu M, Guo L. Multifocus color image fusion using quaternion wavelet transform[C]//Image and Signal Processing (CISP), 2012 5th International Congress on. IEEE, 2012: 543-546.
  6. 6.Bavirisetti D P, Dhuli R. Two-scale image fusion of visible and infrared images using saliency detection[J]. Infrared Physics & Technology, 2016, 76: 52-64.
  7. 7.Li S, Kang X, Hu J. Image fusion with guided filtering[J]. IEEE Transactions on Image Processing, 2013, 22(7): 2864-2875.
  8. 8.Zong J, Qiu T. Medical image fusion based on sparse representation of classified image patches[J]. Biomedical Signal Processing and Control, 2017, 34: 195-205.
  9. 9.Zhang Q, Fu Y, Li H, et al. Dictionary learning method for joint sparse representation-based image fusion[J]. Optical Engineering, 2013, 52(5): 057006.
  10. 10.Gao R, Vorobyov S A, Zhao H. Image fusion with cosparse analysis operator[J]. IEEE Signal Processing Letters, 2017, 24(7): 943-947.
  11. 11.Li H, Wu X J. Multi-focus Image Fusion Using Dictionary Learning and Low-Rank Representation[C]//International Conference on Image and Graphics. Springer, Cham, 2017: 675-686.
  12. 12.Liu Y, Chen X, Ward R K, et al. Image fusion with convolutional sparse representation[J]. IEEE signal processing letters, 2016, 23(12): 1882-1886.
  13. 13.Liu Y, Chen X, Peng H, et al. Multi-focus image fusion with a deep convolutional neural network[J]. Information Fusion, 2017, 36: 191-207.
  14. 14.Huang G, Liu Z, Weinberger K Q, et al. Densely connected convolutional networks[C]//Proceedings of the IEEE conference on computer vision and pattern recognition. 2017, 1(2): 3.
  15. 15.Prabhakar K R, Srikar V S, Babu R V. DeepFuse: A Deep Unsupervised Approach for Exposure Fusion with Extreme Exposure Image Pairs[C]//2017 IEEE International Conference on Computer Vision (ICCV). IEEE, 2017: 4724-4732.
  16. 16.He K, Zhang X, Ren S, et al. Deep residual learning for image recognition[C]//Proceedings of the IEEE conference on computer vision and pattern recognition. 2016: 770-778.
  17. 17.Wang Z, Bovik A C, Sheikh H R, et al. Image quality assessment: from error visibility to structural similarity[J]. IEEE transactions on image processing, 2004, 13(4): 600-612.
  18. 18.Lin T Y, Maire M, Belongie S, et al. Microsoft coco: Common objects in context[C]//European conference on computer vision. Springer, Cham, 2014: 740-755.
  19. 19.Ma J, Zhou Z, Wang B, et al. Infrared and visible image fusion based on visual saliency map and weighted least square optimization[J]. Infrared Physics & Technology, 2017, 82: 8-17.
  20. 20.Toet A.(2014) TNO Image fusion dataset. Figshare. data [online]. https://figshare.com/articles/TN Image Fusion Dataset/1008029.
  21. 21.Li H.(2018) CODE: DenseFuse-A Fusion Approach to Infrared and Visible Image [online]. Avaliable: https://github.com/hli1221/imagefusion_densefuse/blob/master/images
  22. 22.Kumar B K S. Image fusion based on pixel significance using cross bilateral filter[J]. Signal, image and video processing, 2015, 9(5): 1193-1204.
  23. 23.Ma J, Chen C, Li C, et al. Infrared and visible image fusion via gradient transfer and total variation minimization[J]. Information Fusion, 2016, 31: 100-109.
  24. 24.Liu C H, Qi Y, Ding W R. Infrared and visible image fusion method based on saliency detection in sparse domain[J]. Infrared Physics & Technology, 2017, 83: 94-102.
  25. 25.Xydeas C S, Petrovic V. Objective image fusion performance measure[J]. Electronics letters, 2000, 36(4): 308-309.
  26. 26.Aslantas V, Bendes E. A new image quality metric for image fusion: The sum of the correlations of differences[J]. AEU-International Journal of Electronics and Communications, 2015, 69(12): 1890-1896.
  27. 27.Haghighat M, Razian M A. Fast-FMI: non-reference image fusion metric[C]//Application of Information and Communication Technologies (AICT), 2014 IEEE 8th International Conference on. IEEE, 2014: 1-3.
  28. 28.Ma K, Zeng K, Wang Z. Perceptual quality assessment for multi-exposure image fusion[J]. IEEE Transactions on Image Processing, 2015, 24(11): 3345-3356.
  29. 29.Z. Yang, T. Dan, and Y. Yang, Multi-temporal Remote Sensing Image Registration Using Deep Convolutional Features, IEEE Access, vol. 6, pp. 38544-38555, 2018.
  30. 30.M. Zhao, Y. Wu, S. Member, S. Pan, and F. Zhou, Automatic Registration of Images With Inconsistent Content Through Line-Support Region Segmentation and Geometrical Outlier Removal, IEEE Trans. Image Process., vol. 27, no. 6, pp. 2731-2746, 2018.
  31. 31.S. J. Chen, H. L. Shen, C. Li, and J. H. Xin, Normalized Total Gradient: A New Measure for Multispectral Image Registration, IEEE Trans. Image Process., vol. 27, no. 3, pp. 1297-1310, 2017.
  32. 32.Hwang S, Park J, Kim N, et al. Multispectral pedestrian detection: Benchmark dataset and baseline[C]//Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2015: 1037-1045.

Citation

MLA
Li, H., and X.-J. Wu. “DenseFuse: A Fusion Approach to Infrared and Visible Images”. IEEE Transactions on Image Processing, vol. 28, no. 5, 2019, pp. 2614–23, https://doi.org/10.1109/TIP.2018.2887342.
APA
Li, H., & Wu, X.-J. (2019). DenseFuse: A Fusion Approach to Infrared and Visible Images. IEEE Transactions on Image Processing, 28(5), 2614–2623. https://doi.org/10.1109/TIP.2018.2887342
Chicago
Li, H., and X.-J. Wu. 2019. “DenseFuse: A Fusion Approach to Infrared and Visible Images”. IEEE Transactions on Image Processing 28 (5): 2614–23. https://doi.org/10.1109/TIP.2018.2887342.
Harvard
Li, H. and Wu, X.-J. (2019) “DenseFuse: A Fusion Approach to Infrared and Visible Images”, IEEE Transactions on Image Processing, 28(5), pp. 2614–2623. Available at: https://doi.org/10.1109/TIP.2018.2887342.
Vancouver
1. Li H, Wu X-J (2019) DenseFuse: A Fusion Approach to Infrared and Visible Images. IEEE Transactions on Image Processing 28:2614–2623

BibTeX

@article{Li_2019, title={DenseFuse: A Fusion Approach to Infrared and Visible Images}, volume={28}, ISSN={1941-0042}, url={http://dx.doi.org/10.1109/TIP.2018.2887342}, DOI={10.1109/tip.2018.2887342}, number={5}, journal={IEEE Transactions on Image Processing}, publisher={Institute of Electrical and Electronics Engineers (IEEE)}, author={Li, Hui and Wu, Xiao-Jun}, year={2019}, month=May, pages={2614–2623} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF