Image Denoising and Inpainting with Deep Neural Networks

Junyuan XieLinli XuEnhong Chen

article2012NeurIPS1,469 citations

Proposes stacked sparse denoising auto-encoders with a specialized training scheme to perform blind image inpainting and denoising, enabling the automatic removal of complex corruptions such as superimposed text without prior mask information.

Listen

Digital image signals are frequently corrupted by noise from acquisition channels or artificial modifications such as superimposed text. Traditional image restoration methods, including linear sparse coding, often require manual intervention, such as providing exact masks that locate damaged pixels prior to restoration. In real-world operational environments, manually labeling damaged regions is expensive, time-consuming, or impossible. Developing automated methods that can simultaneously locate and repair complex image corruptions without prior knowledge of damaged areas is a critical practical challenge.

The article evaluates a novel image restoration framework combining deep neural networks with sparse representations, termed Stacked Sparse Denoising Auto-encoders. The primary objective is to demonstrate that this deep learning approach can effectively perform image denoising and blind inpainting—automatically identifying and removing complex corruptions without pre-defined defect masks.

To accomplish this, the authors designed a layer-wise pre-training scheme that trains the network to map corrupted image patches directly to their original, clean counterparts. The approach was evaluated using standard natural benchmark images corrupted by varying levels of Gaussian noise and superimposed text of different font sizes. The method's denoising and blind inpainting performance was benchmarked against established linear baseline models, including the widely used K-SVD dictionary-learning algorithm and Gaussian scale-mixture methods.

The experimental results highlight four key findings. First, the deep model achieved Gaussian noise removal performance statistically comparable to established benchmarks, registering signal-to-noise ratios between approximately 24.2 dB and 30.5 dB across heavy to mild noise levels. Second, visual inspections revealed that the deep model produced sharper boundaries and superior texture detail in complex regions compared to traditional alternatives. Third, the model successfully conducted blind inpainting on complex superimposed text, completely erasing small fonts and dimming large fonts without requiring any prior corruption location masks, matching non-blind algorithms. Finally, feature analysis demonstrated that training auto-encoders on realistic, noise-specific distributions yielded higher classification accuracy than training on arbitrary, simple noise distributions.

These findings imply that learned deep representations can overcome the structural limits of shallow linear models in low-level vision tasks. Operationally, the ability to perform blind inpainting removes the cost, delay, and operational friction of manual defect tagging, enabling automated visual preprocessing pipelines. Furthermore, the results indicate that tailoring training data corruption to match real-world operational noise significantly enhances downstream machine learning performance.

Organizations handling specialized image restoration workflows should evaluate data-driven deep networks when defect locations cannot be easily labeled ahead of time. Future development should focus on testing this architecture on related domains highlighted by the source, such as audio denoising, video restoration, image super-resolution, and missing data imputation, as well as optimizing network hyperparameters.

Decision-makers should note that the model is heavily dependent on supervised training and only reliably eliminates noise or corruption patterns present in its training data. While confidence in the evaluated tasks is high based on the empirical results, cautious deployment is warranted when dealing with unstructured or novel noise types that deviate significantly from the training distribution.

Cover for Image Denoising and Inpainting with Deep Neural Networks

Abstract

We present a novel approach to low-level vision problems that combines sparse coding and deep networks pre-trained with denoising auto-encoder (DA). We propose an alternative training scheme that successfully adapts DA, originally designed for unsupervised feature learning, to the tasks of image denoising and blind inpainting. Our method's performance in the image denoising task is comparable to that of KSVD which is a widely used sparse coding technique. More importantly, in blind image inpainting task, the proposed method provides solutions to some complex problems that have not been tackled before. Specifically, we can automatically remove complex patterns like superimposed text from an image, rather than simple patterns like pixels missing at random. Moreover, the proposed method does not need the information regarding the region that requires inpainting to be given a priori. Experimental results demonstrate the effectiveness of the proposed method in the tasks of image denoising and blind inpainting. We also show that our new training scheme for DA is more effective and can improve the performance of unsupervised feature learning.

Table of Contents

  • 1 Introduction
  • 2 Model Description
  • 2.1 Problem Formulation
  • 2.2 Denoising Auto-encoder
  • 2.3 Stacked Sparse Denoising Auto-encoders
  • 3 Experiments
  • 3.1 Denoising White Gaussian Noise
  • 3.2 Image Inpainting
  • 3.3 Hidden Layer Feature Analysis
  • 4 Discussion
  • 4.1 Prior vs. Learned Structure
  • 4.2 Advantages and Limitations
  • 5 Conclusion
  • 6 Acknowledgement
  • References

Knowls

  1. Knowl 1 — Layer-wise Pre-training Scheme for Stacked Sparse Denoising Auto-encoders

    model/method

    Stacked Sparse Denoising Auto-encoders (SSDA) adapt the greedy layer-wise pre-training of standard stacked denoising auto-encoders (SDA) specifically for low-level restoration problems such as image denoising and blind inpainting.

    In standard unsupervised SDA, only clean data yy is provided; noisy inputs η(y)\eta(y) are synthesized on the fly, and activations of subsequent layers are constructed by corrupting the clean hidden representation η(h(y))\eta(h(y)). In contrast, for restoration problems where noisy-clean sample pairs (xi,yi)(x_i, y_i) with xi=η(yi)x_i = \eta(y_i) naturally exist:

    1. The first Denoising Auto-encoder (DA) is trained to map the observed noisy patch xix_i to its corresponding clean patch yiy_i.
    2. For training the (k+1)(k+1)-th DA layer, the hidden activations hk(xi)h_k(x_i) of the noisy input and hk(yi)h_k(y_i) of the clean input produced by the kk-th trained DA layer are directly passed as the noisy and clean training inputs, respectively, rather than applying an artificial corruption operator η(⋅)\eta(\cdot) to the clean hidden representation space.
    3. After pre-training KK such DAs, the weights are unrolled to assemble a feedforward deep neural network containing 1 input layer, 2K−12K-1 hidden layers, and 1 output layer, which is subsequently fine-tuned end-to-end.
  2. Knowl 2 — Optimization Objective for Single-Layer Sparse Denoising Auto-encoder

    equation

    For a training set of NN patch pairs {(xi,yi)}i=1N\{(x_i, y_i)\}_{i=1}^N where xi∈Rnx_i \in \mathbb{R}^n is the corrupted patch and yi∈Rny_i \in \mathbb{R}^n is the clean patch, the parameter set θ={W,b,W′,b′}\theta = \{W, b, W', b'\} of a single sparse denoising auto-encoder layer is learned by minimizing the penalized reconstruction loss:

    L1(X,Y;θ)=1N∑i=1N12∥yi−y^(xi)∥22+βKL(ρ^∥ρ)+λ2(∥W∥F2+∥W′∥F2)L_1(X, Y; \theta) = \frac{1}{N}\sum_{i=1}^N \frac{1}{2}\|y_i - \hat{y}(x_i)\|_2^2 + \beta \text{KL}(\hat{\rho} \parallel \rho) + \frac{\lambda}{2}\left(\|W\|_F^2 + \|W'\|_F^2\right)

    where:

    • h(xi)=σ(Wxi+b)h(x_i) = \sigma(W x_i + b) denotes the hidden layer activations, with σ(z)=(1+exp⁡(−z))−1\sigma(z) = (1 + \exp(-z))^{-1} applied element-wise.
    • y^(xi)=σ(W′h(xi)+b′)\hat{y}(x_i) = \sigma(W' h(x_i) + b') is the reconstructed clean patch.
    • ∥⋅∥F\|\cdot\|_F denotes the Frobenius norm, weighted by weight-decay parameter λ>0\lambda > 0.
    • β>0\beta > 0 controls the sparsity penalty.
    • KL(ρ^∥ρ)=∑j=1∣ρ^∣(ρlog⁡ρρ^j+(1−ρ)log⁡1−ρ1−ρ^j)\text{KL}(\hat{\rho} \parallel \rho) = \sum_{j=1}^{|\hat{\rho}|} \left( \rho \log \frac{\rho}{\hat{\rho}_j} + (1 - \rho) \log \frac{1 - \rho}{1 - \hat{\rho}_j} \right) is the Kullback-Leibler divergence between the desired low target activation ρ\rho (e.g., ρ=0.05\rho = 0.05) and the empirical average activation of hidden unit jj, ρ^j=1N∑i=1Nhj(xi)\hat{\rho}_j = \frac{1}{N}\sum_{i=1}^N h_j(x_i).
  3. Knowl 3 — Global Fine-Tuning Objective for Deep SSDA

    equation

    After pre-training KK sparse denoising auto-encoders, a deep neural network consisting of 2K−12K-1 hidden layers and parameter set θ={Wj,bj}j=12K\theta = \{W_j, b_j\}_{j=1}^{2K} is constructed and fine-tuned on the patch pairs {(xi,yi)}i=1N\{(x_i, y_i)\}_{i=1}^N to minimize:

    L2(X,Y;θ)=1N∑i=1N12∥yi−y(xi)∥22+λ2∑j=12K∥Wj∥F2L_2(X, Y; \theta) = \frac{1}{N}\sum_{i=1}^N \frac{1}{2}\|y_i - y(x_i)\|_2^2 + \frac{\lambda}{2}\sum_{j=1}^{2K} \|W_j\|_F^2

    where y(xi)y(x_i) represents the end-to-end output of the deep network given noisy input xix_i, ∥⋅∥F\|\cdot\|_F denotes the Frobenius norm, and λ\lambda is the weight decay parameter.

    The sparsity regularization term KL(ρ^∥ρ)\text{KL}(\hat{\rho} \parallel \rho) is removed during fine-tuning because the pre-trained weights act as implicit regularization.

  4. Knowl 4 — SSDA Image Restoration Pipeline

    algorithm

    The complete training and inference procedure of the Stacked Sparse Denoising Auto-encoder (SSDA) for image denoising or inpainting operates on overlapping image patches:

    Input: Training clean images, corrupted training images (or noise process η\eta), number of stacked DA stages KK, patch dimensions, hyperparameters β,λ,ρ\beta, \lambda, \rho
    Output: Restored image YoutY_\text{out}
    Extract overlapping patch pairs {(xi,yi)}i=1N\{(x_i, y_i)\}_{i=1}^N from noisy and clean training images
    X(1)={xi}i=1NX^{(1)} = \{x_i\}_{i=1}^N
    Y(1)={yi}i=1NY^{(1)} = \{y_i\}_{i=1}^N
    for k=1k = 1 to KK do
        Train kk-th DA with parameters {Wk,bk,Wk′,bk′}\{W_k, b_k, W'_k, b'_k\} by optimizing single-layer loss L1(X(k),Y(k);θk)L_1(X^{(k)}, Y^{(k)}; \theta_k) via L-BFGS
        X(k+1)={σ(Wkx+bk) for x∈X(k)}X^{(k+1)} = \{\sigma(W_k x + b_k) \text{ for } x \in X^{(k)}\}
        Y(k+1)={σ(Wky+bk) for y∈Y(k)}Y^{(k+1)} = \{\sigma(W_k y + b_k) \text{ for } y \in Y^{(k)}\}
    end for
    Assemble deep network of 2K−12K-1 hidden layers using weights {W1,…,WK,WK′,…,W1′}\{W_1, \dots, W_K, W'_K, \dots, W'_1\}
    Fine-tune all parameters θ\theta to minimize L2(X,Y;θ)L_2(X, Y; \theta) via L-BFGS
    Given unseen noisy test image XtestX_\text{test}:
    Extract all overlapping patches {xj}\{x_j\} from XtestX_\text{test}
    Reconstruct clean patches y^j=y(xj)\hat{y}_j = y(x_j) using the fine-tuned deep network
    Average overlapping patch reconstructions to generate the final restored image YoutY_\text{out}
    return YoutY_\text{out}

    Typical hyperparameters are K=2K=2, hidden layer dimensionality equal to 5 times the input patch dimension, λ=10−4\lambda = 10^{-4}, β=10−2\beta = 10^{-2}, and ρ=0.05\rho = 0.05.

  5. Knowl 5 — Gaussian Noise Image Denoising Performance Comparison

    data/table

    The quantitative image denoising performance of Stacked Sparse Denoising Auto-encoders (SSDA) is evaluated on standard test images corrupted by additive zero-mean white Gaussian noise of varying standard deviation σ∈{25,50,100}\sigma \in \{25, 50, 100\}. Performance is measured by Peak Signal to Noise Ratio (PSNR in dB), defined as PSNR=10log⁡10(2552/σe2)\text{PSNR} = 10 \log_{10}(255^2 / \sigma_e^2), where σe2\sigma_e^2 is the mean squared error.

    Method σ=25\sigma=25 (Noisy PSNR=20.17) σ=50\sigma=50 (Noisy PSNR=14.16) σ=100\sigma=100 (Noisy PSNR=8.13)
    SSDA 30.52±1.0230.52 \pm 1.02 27.37±1.1027.37 \pm 1.10 24.18±1.3924.18 \pm 1.39
    BLS-GSM 30.49±1.1730.49 \pm 1.17 27.28±1.4427.28 \pm 1.44 24.37±1.3624.37 \pm 1.36
    KSVD 30.96±0.7730.96 \pm 0.77 27.34±1.1127.34 \pm 1.11 23.50±1.1523.50 \pm 1.15

    While the overall numerical PSNR differences between SSDA, BLS-GSM, and KSVD are statistically insignificant, SSDA produces sharper edge boundaries and restores more texture details in complex image regions.

  6. Knowl 6 — Blind Inpainting of Complex Superimposed Text

    empirical result

    SSDA is capable of performing blind image inpainting of complex, non-i.i.d. corruption patterns, such as superimposed text with font sizes ranging from 18 to 36 pixels, without requiring an inpainting mask or prior knowledge of corrupted pixel locations.

    SSDA completely removes text of small font sizes and dims text of larger font sizes. When compared against non-blind KSVD inpainting—which is provided with ground-truth masks identifying exactly which pixels require restoration—the blind SSDA achieves comparable qualitative restoration quality despite having to automatically identify and reconstruct corrupted regions simultaneously.

  7. Knowl 7 — Noise-Distribution-Matched Pre-training for Unsupervised Feature Learning

    data/table

    The effectiveness of Denoising Auto-encoders for unsupervised feature learning depends strongly on the correspondence between the pre-training corruption distribution and the real noise distribution in the data. DAs were trained on MNIST digits using three distinct noise types (Gaussian, Salt-and-Pepper, and Image background). The extracted hidden features were then evaluated for handwritten digit classification accuracy using SVMs across all three noise conditions.

    Training Noise Test: Gaussian Test: Salt-and-Pepper Test: Image Background
    Gaussian 91.42% 82.95% 86.45%
    Salt-and-Pepper 90.05% 90.14% 81.77%
    Image Background 84.88% 74.47% 86.87%

    Features learned by a DA pre-trained on a specific noise distribution consistently achieve the highest classification accuracy when evaluated on that same noise type, indicating that realistic and domain-informed noise distributions during DA pre-training produce superior feature representations.

  8. Knowl 8 — Supervised Scope Limitation of SSDA

    limitation

    SSDA derives its restoration capability entirely from supervised training on paired data rather than structural analytical priors (such as wavelet domain Gaussian scale mixtures). While SSDA demonstrates generalization to unseen instances of similar noise characteristics, it is generally limited to removing noise patterns that were represented in the training distribution, making it suited primarily for tasks where the corruption process is known or constrained to a specific family of corruptions.

Coverage note — None was omitted; all contributed models, loss formulations, training algorithms, denoising and inpainting experimental results, feature learning analyses, and stated limitations are covered.

References

  1. 1.J. Xu, K. Zhang, M. Xu, and Z. Zhou. An adaptive threshold method for image denoising based on wavelet domain. Proceedings of SPIE, the International Society for Optical Engineering, 7495:165, 2009.
  2. 2.J. Portilla, V. Strela, M.J. Wainwright, and E.P. Simoncelli. Image denoising using scale mixtures of Gaussians in the wavelet domain. Image Processing, IEEE Transactions on, 12(11):1338–1351, 2003.
  3. 3.F. Luisier, T. Blu, and M. Unser. A new SURE approach to image denoising: Interscale orthonormal wavelet thresholding. IEEE Transactions on Image Processing, 16(3):593–606, 2007.
  4. 4.B.A. Olshausen and D.J. Field. Sparse coding with an overcomplete basis set: A strategy employed by V1? Vision research, 37(23):3311–3325, 1997.
  5. 5.K. Kreutz-Delgado, J.F. Murray, B.D. Rao, K. Engan, T.W. Lee, and T.J. Sejnowski. Dictionary learning algorithms for sparse representation. Neural computation, 15(2):349–396, 2003.
  6. 6.M. Elad and M. Aharon. Image denoising via sparse and redundant representations over learned dictionaries. IEEE Transactions on Image Processing, 15(12):3736–3745, 2006.
  7. 7.J. Mairal, M. Elad, and G. Sapiro. Sparse representation for color image restoration. IEEE Transactions on Image Processing, 17(1):53–69, 2008.
  8. 8.X. Lu, H. Yuan, P. Yan, Y. Yuan, L. Li, and X. Li. Image denoising via improved sparse coding. Proceedings of the British Machine Vision Conference, pages 74–1, 2011.
  9. 9.J. Mairal, F. Bach, J. Ponce, and G. Sapiro. Online dictionary learning for sparse coding. Proceedings of the 26th Annual International Conference on Machine Learning, pages 689–696, 2009.
  10. 10.A. Criminisi, P. Pérez, and K. Toyama. Region filling and object removal by exemplar-based image inpainting. IEEE Transactions on Image Processing, 13(9):1200–1212, 2004.
  11. 11.M. Bertalmio, G. Sapiro, V. Caselles, and C. Ballester. Image inpainting. Proceedings of the 27th annual conference on Computer graphics and interactive techniques, pages 417–424, 2000.
  12. 12.A. Telea. An image inpainting technique based on the fast marching method. Journal of graphics tools., 9(1):23–34, 2004.
  13. 13.B. Dong, H. Ji, J. Li, Z. Shen, and Y. Xu. Wavelet frame based blind image inpainting. Applied and Computational Harmonic Analysis, 2011.
  14. 14.Y. Wang, A. Szlam, and G. Lerman. Robust locally linear analysis with applications to image denoising and blind inpainting. preprint, 2011.
  15. 15.M. Yan. Restoration of images corrupted by impulse noise using blind inpainting and l0 norm. preprint, 2011.
  16. 16.V. Jain and H.S. Seung. Natural image denoising with convolutional networks. Advances in Neural Information Processing Systems, 21:769–776, 2008.
  17. 17.H. Lee, C. Ekanadham, and A. Ng. Sparse deep belief net model for visual area V2. Advances in Neural Information Processing Systems 20, pages 873–880, 2008.
  18. 18.D. Erhan, Y. Bengio, A. Courville, P.A. Manzagol, P. Vincent, and S. Bengio. Why does unsupervised pre-training help deep learning? The Journal of Machine Learning Research, 11:625–660, 2010.
  19. 19.Y. Bengio. Learning deep architectures for AI. Foundations and Trends® in Machine Learning, 2(1):1–127, 2009.
  20. 20.R. Salakhutdinov and G.E. Hinton. Deep boltzmann machines. Proceedings of the international conference on artificial intelligence and statistics, 5(2):448–455, 2009.
  21. 21.P. Vincent, H. Larochelle, I. Lajoie, Y. Bengio, and P.A. Manzagol. Stacked denoising autoencoders: Learning useful representations in a deep network with a local denoising criterion. The Journal of Machine Learning Research, 11:3371–3408, 2010.
  22. 22.Q.V. Le, A. Coates, B. Prochnow, and A.Y. Ng. On optimization methods for deep learning. Learning, pages 265–272, 2011.
  23. 23.S. Roth and M.J. Adviser-Black. High-order markov random fields for low-level vision. Brown University Press, 2007.

Citation

MLA
Xie, J., et al. “Image Denoising and Inpainting with Deep Neural Networks”. Advances in Neural Information Processing Systems, vol. 25, 2012, https://proceedings.neurips.cc/paper_files/paper/2012/file/6cdd60ea0045eb7a6ec44c54d29ed402-Paper.pdf.
APA
Xie, J., Xu, L., & Chen, E. (2012). Image Denoising and Inpainting with Deep Neural Networks. Advances in Neural Information Processing Systems, 25. https://proceedings.neurips.cc/paper_files/paper/2012/file/6cdd60ea0045eb7a6ec44c54d29ed402-Paper.pdf
Chicago
Xie, J., L. Xu, and E. Chen. 2012. “Image Denoising and Inpainting with Deep Neural Networks”. Advances in Neural Information Processing Systems 25. https://proceedings.neurips.cc/paper_files/paper/2012/file/6cdd60ea0045eb7a6ec44c54d29ed402-Paper.pdf.
Harvard
Xie, J., Xu, L. and Chen, E. (2012) “Image Denoising and Inpainting with Deep Neural Networks”, Advances in Neural Information Processing Systems. Curran Associates, Inc. Available at: https://proceedings.neurips.cc/paper_files/paper/2012/file/6cdd60ea0045eb7a6ec44c54d29ed402-Paper.pdf.
Vancouver
1. Xie J, Xu L, Chen E (2012) Image Denoising and Inpainting with Deep Neural Networks. Advances in Neural Information Processing Systems 25:

BibTeX

@inproceedings{xie2012image,
  title = {Image Denoising and Inpainting with Deep Neural Networks},
  author = {Xie, Junyuan and Xu, Linli and Chen, Enhong},
  year = {2012},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {25},
  url = {https://proceedings.neurips.cc/paper_files/paper/2012/file/6cdd60ea0045eb7a6ec44c54d29ed402-Paper.pdf}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors