FFDNet: Toward a Fast and Flexible Solution for CNN-Based Image Denoising

Kai ZhangWangmeng ZuoLei Zhang

article2017IEEE TIP2,678 citationsIEEE Signal Processing Society Best Paper Award

Introduces FFDNet, a convolutional neural network that uses a tunable noise level map to handle both uniform and spatially variant noise across wide noise ranges within a single model while executing faster than BM3D on standard CPUs.

Listen

The article addresses the challenge of removing noise from digital images, a common issue during capture that degrades quality and hinders subsequent computer vision tasks. Existing deep learning approaches typically require separate models for each noise strength and cannot readily handle noise that varies across an image, reducing their practicality.

The article set out to develop and test a single convolutional neural network, called FFDNet, that could manage a wide range of noise levels, spatially varying noise, and real-world noise while remaining fast and effective.

The approach involved training the network on large sets of clean and synthetically noised images, supplying a tunable noise level map as an extra input channel and processing downsampled sub-images to improve speed and receptive field size. Evaluation used standard benchmark datasets for both synthetic additive white Gaussian noise and real photographs, with direct comparisons against established methods such as BM3D, WNNM, and prior CNN models.

FFDNet matched or exceeded the denoising accuracy of leading methods across noise levels from 0 to 75 using one network, delivered roughly three times faster CPU performance than BM3D, and produced visually convincing results on spatially variant and real noise when appropriate noise maps were supplied. It also avoided common artifacts when users selected higher noise levels to trade detail for stronger smoothing.

These outcomes show that a single, flexible model can replace multiple specialized networks, lower computational cost, and support integration into broader image restoration pipelines such as deblurring or super-resolution. The work therefore offers a more deployable solution for practical imaging applications.

Next steps include embedding FFDNet within variable-splitting algorithms for other restoration problems and pairing it with improved noise estimation methods when the exact noise level is unknown. The main limitations are reliance on reasonably accurate noise level inputs and the fact that real-noise performance still benefits from modest user adjustment of the input map; results are robust on controlled synthetic data but should be interpreted with caution for highly complex, uncharacterized real-world noise.

  • Paper: SwinIR: Image Restoration Using Swin Transformer, Jingyun Liang et al. (2021). Extends deep restoration models by replacing pure convolutional backbones like FFDNet with shifted-window self-attention for superior image denoising and super-resolution.
  • Paper: Restormer: Efficient Transformer for High-Resolution Image Restoration, Syed Waqas Zamir et al. (2022). Develops an efficient Transformer-based image restoration architecture that achieves state-of-the-art results on high-resolution synthetic and real-world image denoising.
  • Paper: Multi-Stage Progressive Image Restoration, Syed Waqas Zamir et al. (2021). Proposes a multi-stage progressive restoration framework balancing contextual reasoning and detail preservation across image denoising and deblurring.
  • Paper: Uformer: A General U-Shaped Transformer for Image Restoration, Zhendong Wang et al. (2021). Generalizes multi-scale deep restoration by integrating local window self-attention into a U-Net architecture across diverse image degradation problems.
  • Paper: Pre-Trained Image Processing Transformer, Hanting Chen et al. (2020). Scales universal image restoration through large-scale Transformer pre-training across multiple tasks including denoising, deraining, and super-resolution.
  • Paper: Deep Image Prior, Dmitry Ulyanov et al. (2017). Investigates unsupervised restoration and denoising directly from un-trained generator network architectures without requiring external noisy-clean dataset pairs.
Cover for FFDNet: Toward a Fast and Flexible Solution for CNN-Based Image Denoising

Abstract

Due to the fast inference and good performance, discriminative learning methods have been widely studied in image denoising. However, these methods mostly learn a specific model for each noise level, and require multiple models for denoising images with different noise levels. They also lack flexibility to deal with spatially variant noise, limiting their applications in practical denoising. To address these issues, we present a fast and flexible denoising convolutional neural network, namely FFDNet, with a tunable noise level map as the input. The proposed FFDNet works on downsampled sub-images, achieving a good trade-off between inference speed and denoising performance. In contrast to the existing discriminative denoisers, FFDNet enjoys several desirable properties, including (i) the ability to handle a wide range of noise levels (i.e., [0, 75]) effectively with a single network, (ii) the ability to remove spatially variant noise by specifying a non-uniform noise level map, and (iii) faster speed than benchmark BM3D even on CPU without sacrificing denoising performance. Extensive experiments on synthetic and real noisy images are conducted to evaluate FFDNet in comparison with state-of-the-art denoisers. The results show that FFDNet is effective and efficient, making it highly attractive for practical denoising applications.

Table of Contents

  • I Introduction
  • II Related Work
  • II-A MAP Inference Guided Discriminative Learning
  • II-B Plain Discriminative Learning
  • III Proposed Fast and Flexible Discriminative CNN Denoiser
  • III-A Network Architecture
  • III-B Noise Level Map
  • III-C Denoising on Sub-images
  • III-D Examining the Role of Noise Level Map
  • III-E FFDNet vs. a Single Blind Model
  • III-F Residual vs. Non-residual Learning of Plain CNN
  • III-G Un-clipping vs. Clipping of Noisy Images for Training
  • IV Experiments
  • IV-A Dataset Generation and Network Training
  • IV-B Experiments on AWGN Removal
  • IV-C Experiments on Spatially Variant AWGN Removal
  • IV-D Experiments on Noise Level Sensitivity
  • IV-E Experiments on Real Noisy Images
  • IV-F Running Time
  • V Conclusion
  • References

Knowls

  1. Knowl 1 — FFDNet Architecture and Sub-Image Denoising Pipeline

    model/method

    The Fast and Flexible Denoising Network (FFDNet) performs image denoising directly on downsampled sub-images conditioned on an explicit noise level map.

    Given an observed noisy image yRW×H×Cy \in \mathbb{R}^{W \times H \times C}, where WW is image width, HH is image height, and CC is channel count (C=1C=1 for grayscale, C=3C=3 for RGB color):

    1. Sub-Image Downsampling: A reversible downsampling operator with stride 2 reshapes yy into four sub-images of spatial dimension W2×H2×C\frac{W}{2} \times \frac{H}{2} \times C, producing a concatenated tensor of dimension W2×H2×4C\frac{W}{2} \times \frac{H}{2} \times 4C.

    2. Noise Level Map Concatenation: A tunable noise level map MRW2×H2×1M \in \mathbb{R}^{\frac{W}{2} \times \frac{H}{2} \times 1} is concatenated channel-wise with the downsampled sub-images to form an input tensor y~RW2×H2×(4C+1)\tilde{y} \in \mathbb{R}^{\frac{W}{2} \times \frac{H}{2} \times (4C+1)}.

    3. Convolutional Processing: A feedforward convolutional neural network acts on y~\tilde{y} using a sequence of 3×33 \times 3 convolutions with zero-padding:

      • Layer 1: Conv+ReLU\text{Conv} + \text{ReLU}
      • Middle layers (2 to L1L-1): Conv+Batch Normalization (BN)+ReLU\text{Conv} + \text{Batch Normalization (BN)} + \text{ReLU}
      • Output layer (LL): Conv\text{Conv}
    4. Sub-Pixel Reconstruction: A sub-pixel upscaling operation (the exact inverse of the downsampling operator) transforms the four denoised sub-images of size W2×H2×4C\frac{W}{2} \times \frac{H}{2} \times 4C produced by the final convolution layer back into the full-resolution estimated clean image x^RW×H×C\hat{x} \in \mathbb{R}^{W \times H \times C}.

    Unlike residual learning frameworks that estimate the residual noise image, FFDNet directly predicts the clean image x^\hat{x}.

  2. Knowl 2 — Tunable Noise Level Map Formulation

    model/method

    In classical variational denoising, the estimate x^\hat{x} of a clean image xx from a noisy observation yy corrupted by additive white Gaussian noise (AWGN) with standard deviation σ\sigma is obtained via:

    x^=argminx12σ2yx2+λΦ(x)\hat{x} = \arg\min_x \frac{1}{2\sigma^2} \|y - x\|^2 + \lambda \Phi(x)

    where Φ(x)\Phi(x) is a regularization prior and λ\lambda controls the trade-off between data fidelity and smoothness. By absorbing λ\lambda into σ\sigma, the solution defines an implicit mapping x^=F(y,σ;Θ)\hat{x} = \mathcal{F}(y, \sigma; \Theta).

    To implement this parameterized mapping in a single convolutional network with parameters Θ\Theta, FFDNet addresses the dimensional mismatch between scalar σR\sigma \in \mathbb{R} and image yRW×H×Cy \in \mathbb{R}^{W \times H \times C} by expanding σ\sigma into a spatial noise level map MM.

    For spatially invariant AWGN, MRW×H×1M \in \mathbb{R}^{W \times H \times 1} is a uniform matrix with every entry set to σ\sigma. For spatially variant noise, M(i,j)M(i, j) represents the local noise standard deviation at pixel location (i,j)(i, j). For multidimensional color noise models with zero-mean Gaussian distribution N(0,Σ)\mathcal{N}(0, \Sigma), MM is generalized into a multi-channel covariance map. This formulation allows a single set of network parameters Θ\Theta to handle arbitrary and non-uniform noise levels across the entire input domain.

  3. Knowl 3 — FFDNet Structural Specifications for Grayscale and Color Denoising

    model/method

    FFDNet employs distinct depth and feature width configurations for grayscale and color images to optimize denoising quality and computational efficiency:

    • Grayscale Model (C=1C=1): Comprises L=15L = 15 convolutional layers with 64 feature channels per intermediate layer. Downsampling by a factor of 2 prior to 3×33 \times 3 convolutions expands the effective receptive field to 62×6262 \times 62 pixels on the input image (compared to 31×3131 \times 31 for a standard 15-layer network on full-resolution images).
    • Color Model (C=3C=3): Comprises L=12L = 12 convolutional layers with 96 feature channels per intermediate layer.

    The shallower depth for color denoising encourages the network to exploit inter-channel correlations among the red, green, and blue channels, while expanding the channel width from 64 to 96 provides the representational capacity necessary for multi-channel color inputs, yielding an average performance gain of approximately 0.15 dB0.15\text{ dB} PSNR over narrower configurations.

  4. Knowl 4 — Mean Squared Error Training Objective for FFDNet

    equation

    FFDNet network parameters Θ\Theta are optimized over a training set of NN degraded-clean image patch pairs {(yi,Mi;xi)}i=1N\{(y_i, M_i; x_i)\}_{i=1}^N by minimizing the mean squared error loss function:

    L(Θ)=12Ni=1NF(yi,Mi;Θ)xi2\mathcal{L}(\Theta) = \frac{1}{2N} \sum_{i=1}^N \|\mathcal{F}(y_i, M_i; \Theta) - x_i\|^2

    where xix_i denotes the ii-th ground-truth clean image patch, yiy_i is the corresponding noisy patch generated by adding noise to xix_i, MiM_i is the associated noise level map, and F(yi,Mi;Θ)\mathcal{F}(y_i, M_i; \Theta) denotes the output of FFDNet.

  5. Knowl 5 — FFDNet Training Dataset and Optimization Protocol

    experimental setup

    The FFDNet training set is composed of 400 images from the Berkeley Segmentation Dataset (BSD), 400 validation images from ImageNet, and 4,744 images from the Waterloo Exploration Database.

    • Patch Sampling: In each training epoch, N=128×8,000=1,024,000N = 128 \times 8,000 = 1,024,000 patches are randomly cropped. Patch sizes are 70×7070 \times 70 for grayscale images and 50×5050 \times 50 for color images to exceed the network receptive field.
    • Noise Perturbation: Additive white Gaussian noise with standard deviation σ\sigma uniformly sampled from [0,75][0, 75] is added to clean patches without 8-bit integer quantization.
    • Optimization: The network is trained using the Adam optimizer with a mini-batch size of 128 and data augmentation via random rotations and horizontal/vertical flips.
    • Learning Rate Schedule: Training begins with a learning rate of 10310^{-3}, which decays to 10410^{-4} when training loss ceases to decrease. Once loss stabilizes for five consecutive epochs, the parameters of each Batch Normalization layer are merged into the adjacent convolution filters. The network is then fine-tuned for an additional 50 epochs at a learning rate of 10610^{-6}.
  6. Knowl 6 — Orthogonal Filter Initialization for Noise-Detail Trade-Off Control

    model/method

    When a non-blind denoiser is conditioned on an explicit noise level map MM, setting the input noise level σ\sigma higher than the true noise level in the image is often desired to suppress texture details and smooth flat areas. However, mismatched input noise maps can induce severe structured visual artifacts in deep networks.

    To ensure that MM robustly controls the trade-off between noise removal and detail preservation without introducing visual artifacts, FFDNet employs orthogonal initialization on the convolutional filters. Orthogonal initialization regularizes filter correlations, enhances gradient propagation during backpropagation, improves network compactness, and suppresses artifact generation under noise level overestimation.

  7. Knowl 7 — Grayscale Denoising Performance on BSD68 and Set12 Benchmarks

    data/table

    The table below summarizes average Peak Signal-to-Noise Ratio (PSNR in dB) for grayscale image denoising on the BSD68 and Set12 datasets across AWGN noise levels σ{15,25,35,50,75}\sigma \in \{15, 25, 35, 50, 75\}, comparing FFDNet with model-based methods (BM3D, WNNM) and discriminative models (MLP, TNRD, DnCNN):

    Dataset Method σ=15\sigma = 15 σ=25\sigma = 25 σ=35\sigma = 35 σ=50\sigma = 50 σ=75\sigma = 75
    BSD68 BM3D 31.07 28.57 27.08 25.62 24.21
    WNNM 31.37 28.83 27.30 25.87 24.40
    MLP 28.96 27.50 26.03 24.59
    TNRD 31.42 28.92 25.97
    DnCNN 31.72 29.23 27.69 26.23 24.64
    FFDNet 31.63 29.19 27.73 26.29 24.79
    Set12 BM3D 32.37 29.97 28.40 26.72 24.91
    WNNM 32.70 30.26 28.69 27.05 25.23
    MLP 30.03 28.46 26.78 25.07
    TNRD 32.50 30.06 26.81
    DnCNN 32.86 30.43 28.82 27.18 25.20
    FFDNet 32.75 30.43 28.92 27.32 25.49

    FFDNet outperforms BM3D by a large margin across all noise levels and exceeds WNNM, MLP, and TNRD by approximately 0.2 dB0.2\text{ dB} on BSD68. While DnCNN achieves slightly higher PSNR at lower noise levels (σ25\sigma \le 25), FFDNet surpasses DnCNN at higher noise levels (σ>25\sigma > 25) due to its larger receptive field.

  8. Knowl 8 — Color Denoising Performance on CBSD68, Kodak24, and McMaster Benchmarks

    data/table

    The table below reports average PSNR (in dB) on color benchmark datasets (CBSD68, Kodak24, McMaster) across noise levels σ{15,25,35,50,75}\sigma \in \{15, 25, 35, 50, 75\} for CBM3D, CDnCNN, and FFDNet:

    Dataset Method σ=15\sigma = 15 σ=25\sigma = 25 σ=35\sigma = 35 σ=50\sigma = 50 σ=75\sigma = 75
    CBSD68 CBM3D 33.52 30.71 28.89 27.38 25.74
    CDnCNN 33.89 31.23 29.58 27.92 24.47
    FFDNet 33.87 31.21 29.58 27.96 26.24
    Kodak24 CBM3D 34.28 31.68 29.90 28.46 26.82
    CDnCNN 34.48 32.03 30.46 28.85 25.04
    FFDNet 34.63 32.13 30.57 28.98 27.27
    McMaster CBM3D 34.06 31.66 29.92 28.51 26.79
    CDnCNN 33.44 31.51 30.14 28.61 25.10
    FFDNet 34.66 32.35 30.81 29.18 27.33

    FFDNet consistently outperforms CBM3D across all test sets and noise levels. It matches or exceeds CDnCNN at low-to-moderate noise levels and substantially outperforms CDnCNN at σ=75\sigma = 75 (e.g., 26.24 dB26.24\text{ dB} vs. 24.47 dB24.47\text{ dB} on CBSD68, and 27.33 dB27.33\text{ dB} vs. 25.10 dB25.10\text{ dB} on McMaster).

  9. Knowl 9 — FFDNet Inference Runtime and Efficiency

    data/table

    The execution runtime (in seconds) of BM3D, DnCNN, and FFDNet evaluated on an Intel Core i7-5820K CPU @ 3.3GHz and an Nvidia Titan X Pascal GPU across image sizes 256×256256 \times 256, 512×512512 \times 512, and 1024×10241024 \times 1024:

    Method Device 256×256256 \times 256 512×512512 \times 512 1024×10241024 \times 1024
    Gray Color Gray Color Gray Color
    BM3D CPU (ST) 0.59 0.98 2.52 3.57 10.77 20.15
    DnCNN CPU (ST) 2.14 2.44 8.63 9.85 32.82 38.11
    CPU (MT) 0.74 0.98 3.41 4.10 12.10 15.48
    GPU 0.011 0.014 0.033 0.040 0.124 0.167
    FFDNet CPU (ST) 0.44 0.62 1.81 2.51 7.24 10.17
    CPU (MT) 0.18 0.21 0.73 0.98 2.96 3.95
    GPU 0.006 0.008 0.012 0.017 0.038 0.057

    Here CPU (ST) denotes single-threaded CPU execution and CPU (MT) denotes multi-threaded CPU execution. On CPU with multi-threading, FFDNet is roughly 3×3\times faster than both BM3D and DnCNN. On GPU, FFDNet denoises a 512×512512 \times 512 color image in 17 milliseconds and a 1024×10241024 \times 1024 image in 57 milliseconds.

  10. Knowl 10 — Effect of Sub-Image Downsampling on Receptive Field and Computational Cost

    empirical result

    Operating on downsampled sub-images provides two primary advantages over full-resolution CNN denoising:

    1. Receptive Field Expansion: A 15-layer network using 3×33 \times 3 convolutions on full-resolution images has a receptive field of 31×3131 \times 31. By performing convolutions on sub-images downsampled by a factor of 2, the effective receptive field on the original image expands to 62×6262 \times 62 without increasing depth or using dilated convolutions (which tend to generate artifacts around sharp edges).
    2. Accuracy and Speed Trade-off: Compared against a baseline 15-layer CNN operating without downsampling on the BSD68 dataset:
      • At low noise (σ=15\sigma = 15), the full-resolution baseline slightly outperforms FFDNet by 0.02 dB0.02\text{ dB} PSNR.
      • At high noise (σ=50\sigma = 50), FFDNet outperforms the full-resolution baseline by 0.09 dB0.09\text{ dB} PSNR due to its larger receptive field.
      • FFDNet executes approximately 3×3\times faster and requires substantially less memory during inference.
  11. Knowl 11 — Non-Residual Learning Dynamics with Batch Normalization

    empirical result

    When plain CNNs are trained for AWGN denoising at a single noise level, predicting residual noise rather than the clean image accelerates training because the residual target follows a zero-mean Gaussian distribution matching Batch Normalization (BN) assumptions.

    However, when training across a wide noise range (σ[0,75]\sigma \in [0, 75]) with a conditioning noise level map MM:

    • Batch normalization consistently accelerates training convergence for both residual and non-residual learning.
    • Although residual learning converges faster in initial training epochs, the final denoising performance after learning rate decay and fine-tuning is identical between residual and direct non-residual formulations.
    • Consequently, moderately deep networks (L<20L < 20) do not require residual skip connections for effective multi-level denoising.
  12. Knowl 12 — Denoising Performance Under 8-Bit Quantization and Clipping

    data/table

    When noisy images are clipped to the standard [0,255][0, 255] integer range (8-bit quantization), noise at extreme pixel values becomes non-zero-mean. The model variant trained specifically on clipped noisy data, FFDNet-Clip, was evaluated on the Clip300 dataset (100 images from BSD300 test set and 200 images from PASCAL VOC 2012) across noise levels σ{15,25,35,50,60}\sigma \in \{15, 25, 35, 50, 60\}:

    Method σ=15\sigma = 15 σ=25\sigma = 25 σ=35\sigma = 35 σ=50\sigma = 50 σ=60\sigma = 60
    DCGRF 31.35 28.67 27.08 25.38 24.45
    RBDN 31.05 28.77 27.31 25.80 23.25
    FFDNet-Clip 31.68 29.25 27.75 26.25 25.51

    FFDNet-Clip outperforms Deep Gaussian Conditional Random Field (DCGRF) and Recursively Branched Deconvolutional Network (RBDN) across all evaluated noise levels on clipped images. Furthermore, FFDNet models trained on unquantized data remain effective when applied directly to 8-bit quantized real noisy images.

Coverage note — None was omitted; all key contributions including architecture design, noise map conditioning, filter orthogonalization, training protocol, ablation studies, benchmark evaluations, clipping experiments, and runtime metrics are fully represented.

References

  1. 1.H. C. Andrews and B. R. Hunt, “Digital image restoration,” Prentice-Hall Signal Processing Series, Englewood Cliffs: Prentice-Hall, 1977, vol. 1, 1977.
  2. 2.P. Chatterjee and P. Milanfar, “Is denoising dead?” IEEE Transactions on Image Processing, vol. 19, no. 4, pp. 895–911, 2010.
  3. 3.S. Roth and M. J. Black, “Fields of experts: A framework for learning image priors,” in IEEE Computer Society Conference on Computer Vision and Pattern Recognition, vol. 2, 2005, pp. 860–867.
  4. 4.D. Zoran and Y. Weiss, “From learning models of natural image patches to whole image restoration,” in IEEE International Conference on Computer Vision, 2011, pp. 479–486.
  5. 5.S. Gu, L. Zhang, W. Zuo, and X. Feng, “Weighted nuclear norm minimization with application to image denoising,” in IEEE Conference on Computer Vision and Pattern Recognition, 2014, pp. 2862–2869.
  6. 6.M. V. Afonso, J. M. Bioucas-Dias, and M. A. Figueiredo, “Fast image recovery using variable splitting and constrained optimization,” IEEE Transactions on Image Processing, vol. 19, no. 9, pp. 2345–2356, 2010.
  7. 7.F. Heide, M. Steinberger, Y.-T. Tsai, M. Rouf, D. Pajak, D. Reddy, O. Gallo, J. Liu, W. Heidrich, K. Egiazarian et al., “FlexISP: A flexible camera image processing framework,” ACM Transactions on Graphics, vol. 33, no. 6, p. 231, 2014.
  8. 8.Y. Romano, M. Elad, and P. Milanfar, “The little engine that could: Regularization by denoising (RED),” submitted to SIAM Journal on Imaging Sciences, 2016.
  9. 9.K. Zhang, W. Zuo, S. Gu, and L. Zhang, “Learning deep CNN denoiser prior for image restoration,” in IEEE Conference on Computer Vision and Pattern Recognition, 2017, pp. 3929–3938.
  10. 10.J. Portilla, V. Strela, M. J. Wainwright, and E. P. Simoncelli, “Image denoising using scale mixtures of gaussians in the wavelet domain,” IEEE Transactions on Image processing, vol. 12, no. 11, pp. 1338–1351, 2003.
  11. 11.K. Dabov, A. Foi, V. Katkovnik, and K. Egiazarian, “Image denoising by sparse 3-D transform-domain collaborative filtering,” IEEE Transactions on Image Processing, vol. 16, no. 8, pp. 2080–2095, 2007.
  12. 12.J. Mairal, F. Bach, J. Ponce, G. Sapiro, and A. Zisserman, “Non-local sparse models for image restoration,” in IEEE International Conference on Computer Vision, 2009, pp. 2272–2279.
  13. 13.W. Dong, L. Zhang, G. Shi, and X. Li, “Nonlocally centralized sparse representation for image restoration,” IEEE Transactions on Image Processing, vol. 22, no. 4, pp. 1620–1630, 2013.
  14. 14.M. Elad and M. Aharon, “Image denoising via sparse and redundant representations over learned dictionaries,” IEEE Transactions on Image Processing, vol. 15, no. 12, pp. 3736–3745, 2006.
  15. 15.J. Mairal, M. Elad, and G. Sapiro, “Sparse representation for color image restoration,” IEEE Transactions on Image Processing, vol. 17, no. 1, pp. 53–69, 2008.
  16. 16.A. Buades, B. Coll, and J.-M. Morel, “A non-local algorithm for image denoising,” in IEEE Conference on Computer Vision and Pattern Recognition, vol. 2, 2005, pp. 60–65.
  17. 17.Y. Chen and T. Pock, “Trainable nonlinear reaction diffusion: A flexible framework for fast and effective image restoration,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 39, no. 6, pp. 1256–1272, 2017.
  18. 18.H. C. Burger, C. J. Schuler, and S. Harmeling, “Image denoising: Can plain neural networks compete with BM3D?” in IEEE Conference on Computer Vision and Pattern Recognition, 2012, pp. 2392–2399.
  19. 19.V. Jain and S. Seung, “Natural image denoising with convolutional networks,” in Advances in Neural Information Processing Systems, 2009, pp. 769–776.
  20. 20.K. Zhang, W. Zuo, Y. Chen, D. Meng, and L. Zhang, “Beyond a Gaussian denoiser: Residual learning of deep CNN for image denoising,” IEEE Transactions on Image Processing, vol. 26, no. 7, pp. 3142–3155, July 2017.
  21. 21.A. Barbu, “Training an active random field for real-time image denoising,” IEEE Transactions on Image Processing, vol. 18, no. 11, pp. 2451–2462, 2009.
  22. 22.K. G. Samuel and M. F. Tappen, “Learning optimized MAP estimates in continuously-valued MRF models,” in IEEE Conference on Computer Vision and Pattern Recognition, 2009, pp. 477–484.
  23. 23.J. Sun and M. F. Tappen, “Learning non-local range markov random field for image restoration,” in IEEE Conference on Computer Vision and Pattern Recognition, 2011, pp. 2745–2752.
  24. 24.U. Schmidt and S. Roth, “Shrinkage fields for effective image restoration,” in IEEE Conference on Computer Vision and Pattern Recognition, 2014, pp. 2774–2781.
  25. 25.U. Schmidt, “Half-quadratic inference and learning for natural images,” Ph.D. dissertation, Technische Universitat, Darmstadt, 2017. [Online]. Available: http://tuprints.ulb.tu-darmstadt.de/6044/
  26. 26.S. Lefkimmiatis, “Non-local color image denoising with convolutional neural networks,” in IEEE Conference on Computer Vision and Pattern Recognition, 2017, pp. 3587–3596.
  27. 27.P. Qiao, Y. Dou, W. Feng, R. Li, and Y. Chen, “Learning non-local image diffusion for image denoising,” in Proceedings of the 2017 ACM on Multimedia Conference, 2017, pp. 1847–1855.
  28. 28.R. Vemulapalli, O. Tuzel, and M.-Y. Liu, “Deep gaussian conditional random field network: A model-based deep network for discriminative denoising,” in IEEE Conference on Computer Vision and Pattern Recognition, June 2016.
  29. 29.J. Kruse, C. Rother, and U. Schmidt, “Learning to push the limits of efficient FFT-based image deconvolution,” in IEEE International Conference on Computer Vision, Oct 2017.
  30. 30.J. Xie, L. Xu, and E. Chen, “Image denoising and inpainting with deep neural networks,” in Advances in Neural Information Processing Systems, 2012, pp. 341–349.
  31. 31.F. Agostinelli, M. R. Anderson, and H. Lee, “Robust image denoising with multi-column deep neural networks,” in Advances in Neural Information Processing Systems, 2013, pp. 1493–1501.
  32. 32.S. Ioffe and C. Szegedy, “Batch normalization: Accelerating deep network training by reducing internal covariate shift,” in International Conference on Machine Learning, 2015, pp. 448–456.
  33. 33.F. Yu and V. Koltun, “Multi-scale context aggregation by dilated convolutions,” in International Conference on Learning Representations, 2016.
  34. 34.X. Mao, C. Shen, and Y.-B. Yang, “Image restoration using very deep convolutional encoder-decoder networks with symmetric skip connections,” in Advances in Neural Information Processing Systems, 2016, pp. 2802–2810.
  35. 35.V. Santhanam, V. I. Morariu, and L. S. Davis, “Generalized deep image to image regression,” in IEEE Conference on Computer Vision and Pattern Recognition, 2017, pp. 5609–5619.
  36. 36.Y. Tai, J. Yang, X. Liu, and C. Xu, “Memnet: A persistent memory network for image restoration,” in IEEE International Conference on Computer Vision, 2017, pp. 4539–4547.
  37. 37.A. Krizhevsky, I. Sutskever, and G. E. Hinton, “ImageNet classification with deep convolutional neural networks,” in Advances in neural information processing systems, 2012, pp. 1097–1105.
  38. 38.S. Nam, Y. Hwang, Y. Matsushita, and S. Joo Kim, “A holistic approach to cross-channel image noise modeling and its application to image denoising,” in IEEE Conference on Computer Vision and Pattern Recognition, June 2016.
  39. 39.W. Shi, J. Caballero, F. Huszar, J. Totz, A. P. Aitken, R. Bishop, D. Rueckert, and Z. Wang, “Real-time single image and video super-resolution using an efficient sub-pixel convolutional neural network,” in IEEE Conference on Computer Vision and Pattern Recognition, 2016, pp. 1874–1883.
  40. 40.A. Levin and B. Nadler, “Natural image denoising: Optimality and inherent bounds,” in IEEE Conference on Computer Vision and Pattern Recognition, 2011, pp. 2833–2840.
  41. 41.D. Wang, P. Cui, M. Ou, and W. Zhu, “Deep multimodal hashing with orthogonal regularization.” in International Joint Conference on Artificial Intelligence, 2015, pp. 2291–2297.
  42. 42.Z. Mhammedi, A. Hellicar, A. Rahman, and J. Bailey, “Efficient orthogonal parametrisation of recurrent neural networks using householder reflections,” arXiv preprint arXiv:1612.00188, 2016.
  43. 43.K. Jia, “Improving training of deep neural networks via singular value bounding,” in IEEE Conference on Computer Vision and Pattern Recognition, 2017, pp. 4344–4352.
  44. 44.D. Xie, J. Xiong, and S. Pu, “All you need is beyond a good init: Exploring better solution for training extremely deep convolutional neural networks with orthonormality and modulation,” in IEEE Conference on Computer Vision and Pattern Recognition, 2017, pp. 6176–6185.
  45. 45.Y. Sun, L. Zheng, W. Deng, and S. Wang, “SVDNet for pedestrian retrieval,” arXiv preprint arXiv:1703.05693, 2017.
  46. 46.D. Mishkin and J. Matas, “All you need is a good init,” ArXiv e-prints, 2015.
  47. 47.G. Riegler, S. Schulter, M. Ruther, and H. Bischof, “Conditioned regression models for non-blind single image super-resolution,” in IEEE International Conference on Computer Vision, 2015, pp. 522–530.
  48. 48.S. H. Chan, X. Wang, and O. A. Elgendy, “Plug-and-Play ADMM for image restoration: Fixed-point convergence and applications,” IEEE Transactions on Computational Imaging, vol. 3, no. 1, pp. 84–98, 2017.
  49. 49.S. Zagoruyko and N. Komodakis, “Diracnets: training very deep neural networks without skip-connections,” arXiv preprint arXiv:1706.00388, 2017.
  50. 50.D. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in International Conference for Learning Representations, 2015.
  51. 51.T. Plotz and S. Roth, “Benchmarking denoising algorithms with real photographs,” in The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), July 2017.
  52. 52.J.-S. Lee, “Refined filtering of image noise using local statistics,” Computer graphics and image processing, vol. 15, no. 4, pp. 380–389, 1981.
  53. 53.J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “ImageNet: A large-scale hierarchical image database,” in IEEE Conference on Computer Vision and Pattern Recognition, 2009, pp. 248–255.
  54. 54.K. Ma, Z. Duanmu, Q. Wu, Z. Wang, H. Yong, H. Li, and L. Zhang, “Waterloo exploration database: New challenges for image quality assessment models,” IEEE Transactions on Image Processing, vol. 26, no. 2, pp. 1004–1016, 2017.
  55. 55.A. Vedaldi and K. Lenc, “MatConvNet: Convolutional neural networks for matlab,” in ACM Conference on Multimedia Conference, 2015, pp. 689–692.
  56. 56.M. Lebrun, M. Colom, and J.-M. Morel, “The noise clinic: A blind image denoising algorithm,” Image Processing On Line, vol. 5, pp. 1–54, 2015. [Online]. Available: http://demo.ipol.im/demo/125/
  57. 57.D. Martin, C. Fowlkes, D. Tal, and J. Malik, “A database of human segmented natural images and its application to evaluating segmentation algorithms and measuring ecological statistics,” in Proc. 8th Int’l Conf. Computer Vision, vol. 2, July 2001, pp. 416–423.
  58. 58.M. Everingham, S. M. A. Eslami, L. Van Gool, C. K. I. Williams, J. Winn, and A. Zisserman, “The pascal visual object classes challenge: A retrospective,” International Journal of Computer Vision, vol. 111, no. 1, pp. 98–136, Jan 2015.
  59. 59.R. Franzen, “Kodak lossless true color image suite,” source: http://r0k.us/graphics/kodak, vol. 4, 1999.
  60. 60.L. Zhang, X. Wu, A. Buades, and X. Li, “Color demosaicking by local directional interpolation and nonlocal adaptive thresholding,” Journal of Electronic Imaging, vol. 20, no. 2, pp. 1–15, 2011.
  61. 61.[Online]. Available: https://ni.neatvideo.com/home
  62. 62.C. Liu, R. Szeliski, S. B. Kang, C. L. Zitnick, and W. T. Freeman, “Automatic estimation and removal of noise from a single image,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 30, no. 2, pp. 299–314, 2008.
  63. 63.M. Colom, M. Lebrun, A. Buades, and J.-M. Morel, “A non-parametric approach for the estimation of intensity-frequency dependent noise,” in IEEE International Conference on Image Processing, 2014, pp. 4261–4265.
  64. 64.L. Azzari and A. Foi, “Gaussian-cauchy mixture modeling for robust signal-dependent noise estimation,” in 2014 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2014, pp. 5357–5361.

Citation

MLA
Zhang, K., et al. “FFDNet: Toward a Fast and Flexible Solution for CNN-Based Image Denoising”. IEEE Transactions on Image Processing, vol. 27, no. 9, 2018, pp. 4608–22, https://doi.org/10.1109/TIP.2018.2839891.
APA
Zhang, K., Zuo, W., & Zhang, L. (2018). FFDNet: Toward a Fast and Flexible Solution for CNN-Based Image Denoising. IEEE Transactions on Image Processing, 27(9), 4608–4622. https://doi.org/10.1109/TIP.2018.2839891
Chicago
Zhang, K., W. Zuo, and L. Zhang. 2018. “FFDNet: Toward a Fast and Flexible Solution for CNN-Based Image Denoising”. IEEE Transactions on Image Processing 27 (9): 4608–22. https://doi.org/10.1109/TIP.2018.2839891.
Harvard
Zhang, K., Zuo, W. and Zhang, L. (2018) “FFDNet: Toward a Fast and Flexible Solution for CNN-Based Image Denoising”, IEEE Transactions on Image Processing, 27(9), pp. 4608–4622. Available at: https://doi.org/10.1109/TIP.2018.2839891.
Vancouver
1. Zhang K, Zuo W, Zhang L (2018) FFDNet: Toward a Fast and Flexible Solution for CNN-Based Image Denoising. IEEE Transactions on Image Processing 27:4608–4622

BibTeX

@article{Zhang_2018, title={FFDNet: Toward a Fast and Flexible Solution for CNN-Based Image Denoising}, volume={27}, ISSN={1941-0042}, url={http://dx.doi.org/10.1109/TIP.2018.2839891}, DOI={10.1109/tip.2018.2839891}, number={9}, journal={IEEE Transactions on Image Processing}, publisher={Institute of Electrical and Electronics Engineers (IEEE)}, author={Zhang, Kai and Zuo, Wangmeng and Zhang, Lei}, year={2018}, month=Sept, pages={4608–4622} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF