Non-local sparse models for image restoration

Julien MairalFrancis BachJean PonceGuillermo SapiroAndrew Zisserman

article2009ICCV1,859 citations

Unifies non-local image self-similarity with learned dictionary sparse coding via simultaneous sparse approximations, achieving state-of-the-art performance in camera raw image denoising and demosaicking at practical computational costs.

Listen

Consumer and mobile digital cameras frequently produce noisy raw image data when operating under low light or high shutter speeds. Reconstructing clean full-color images from this raw sensor data requires solving two core restoration problems: removing sensor noise and reconstructing full color information from incomplete sensor measurements. Prior techniques have addressed these issues either by exploiting image self-similarity to average out noise or by using learned sparse coding to approximate local image patches with a compact set of basis elements. However, each approach presents trade-offs: self-similarity methods struggle when unique image patches lack duplicates, whereas sparse coding can introduce visual artifacts because mathematically similar patches may be reconstructed using inconsistent sets of basis elements.

The article develops and evaluates a unified restoration framework called learned simultaneous sparse coding. This method forces clusters of similar image patches to share identical active elements from a learned dictionary, effectively merging non-local self-similarity with adaptive sparse representations.

To test this framework, the authors conducted quantitative benchmark experiments and qualitative assessments across image denoising and color reconstruction tasks. Denoising was tested against synthetic Gaussian noise across twelve standard test images at multiple noise levels, comparing results with leading methods such as block matching 3D filtering and previous sparse models. Color reconstruction was evaluated on a twenty-four-image benchmark dataset. To establish practical feasibility, the method was also applied directly to noisy raw photographs captured with a consumer digital camera under challenging high-sensitivity settings.

The experimental findings show that the unified framework consistently outperforms or matches established state-of-the-art baselines. First, in synthetic noise reduction, the method produced the highest overall image quality scores, outperforming leading algorithms particularly in heavy noise conditions. Second, in color reconstruction benchmarks, the approach achieved an average quality improvement of approximately 0.87 decibels over top-performing specialized baselines while eliminating visible color distortion artifacts that occur in standard sparse coding. Third, qualitative tests on raw high-sensitivity camera captures demonstrated that the framework produced cleaner results with fewer edge and color artifacts than several established commercial photographic processing tools, even when operating without camera-specific sensor noise models.

These results show that combining non-local patch grouping with dictionary learning significantly improves visual restoration quality without requiring application-specific hand-tuning. The framework offers an effective, unified mechanism for processing raw sensor data into high-quality images. The authors note that the method can be tuned flexibly: adjusting dictionary size and iteration depth allows processing time to drop by an order of magnitude (from around twenty seconds down to sub-second or low-second execution per image) with virtually no noticeable drop in visual fidelity, making it viable for practical software and hardware imaging pipelines.

For operational implementation, system designers and engineering teams should consider adopting simultaneous sparse coding as a general-purpose processing layer for computational imaging pipelines. Prior to deployment, developers should incorporate device-specific, non-uniform noise models to improve handling of non-homogeneous sensor noise and background artifacts. The authors also recommend extending the framework to related restoration challenges, including image deblurring, missing-region completion, and video restoration.

The main limitation noted in the article is the reliance on a uniform noise assumption during real-world camera testing, which occasionally left faint residual noise artifacts in complex backgrounds. Computational cost is also higher than fixed-basis transforms, requiring cluster-based approximations to achieve practical speeds. Despite these limitations, there is high confidence in the quantitative and visual gains demonstrated across standardized image benchmarks.

  • Paper: A non-local algorithm for image denoising, Antoni Buades et al. (2005). Read this first to understand the non-local means principle of grouping similar image neighborhoods that the source combines with sparse coding.
  • Paper: Online dictionary learning for sparse coding, Julien Mairal et al. (2009). Its online dictionary-learning method supplies the learned-dictionary foundation needed to follow the source’s simultaneous sparse-coding model.
Cover for Non-local sparse models for image restoration

Abstract

We propose in this paper to unify two different approaches to image restoration: On the one hand, learning a basis set (dictionary) adapted to sparse signal descriptions has proven to be very effective in image reconstruction and classification tasks. On the other hand, explicitly exploiting the self-similarities of natural images has led to the successful non-local means approach to image restoration. We propose simultaneous sparse coding as a framework for combining these two approaches in a natural manner. This is achieved by jointly decomposing groups of similar signals on subsets of the learned dictionary. Experimental results in image denoising and demosaicking tasks with synthetic and real noise show that the proposed method outperforms the state of the art, making it possible to effectively restore raw images from digital cameras at a reasonable speed and memory cost.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 2.1. Non-Local Means Filtering
  • 2.2. Learned Sparse Coding
  • 2.3. Block Matching 3D (BM3D)
  • 3. Proposed Formulation
  • 3.1. Simultaneous Sparse Coding
  • 3.2. Principle of the Formulation
  • 3.3. Practical Formulation and Implementation
  • 3.4. Real Images and Demosaicking
  • 4. Experimental Validation
  • 4.1. Denoising - Synthetic Noise
  • 4.2. Demosaicking
  • 4.3. Denoising - Real Noise
  • 5. Conclusion
  • Acknowledgments
  • References

Knowls

  1. Knowl 1 — Learned Simultaneous Sparse Coding (LSSC) Formulation

    model/method

    Learned Simultaneous Sparse Coding (LSSC) integrates non-local self-similarities into learned dictionary sparse representations by forcing groups of similar patches to share the same dictionary elements in their sparse decompositions.

    Let a noisy image of nn pixels be represented as overlapping patches yi∈Rmy_i \in \mathbb{R}^m of size mm centered at each pixel i∈{1,…,n}i \in \{1, \dots, n\}. For each patch yiy_i, let Si⊆{1,…,n}S_i \subseteq \{1, \dots, n\} denote the set of indices of patches similar to yiy_i. Let D∈Rm×kD \in \mathbb{R}^{m \times k} be a dictionary whose columns (atoms) belong to the unit ℓ2\ell_2-sphere C={d∈Rm∣∥d∥2=1}\mathcal{C} = \{d \in \mathbb{R}^m \mid \|d\|_2 = 1\}, and let Ai=[αij]j∈Si∈Rk×∣Si∣A_i = [\alpha_{ij}]_{j \in S_i} \in \mathbb{R}^{k \times |S_i|} be the matrix of sparse codes for the patch group SiS_i, where αij∈Rk\alpha_{ij} \in \mathbb{R}^k represents the decomposition vector of patch yjy_j.

    The dictionary DD and sparse code matrices (Ai)i=1n(A_i)_{i=1}^n are learned jointly by solving the optimization problem:

    min⁡(Ai)i=1n,D∈C∑i=1n∥Ai∥p,q∣Si∣ps.t.∀i,∑j∈Si∥yj−Dαij∥22≤εi\min_{(A_i)_{i=1}^n, D \in \mathcal{C}} \sum_{i=1}^n \frac{\|A_i\|_{p,q}}{|S_i|^p} \quad \text{s.t.} \quad \forall i, \sum_{j \in S_i} \|y_j - D\alpha_{ij}\|_2^2 \le \varepsilon_i

    where ∥Ai∥p,q≜∑l=1k∥αil∥pq\|A_i\|_{p,q} \triangleq \sum_{l=1}^k \|\alpha_i^l\|_p^q is the grouped-sparsity (pseudo) matrix norm with αil\alpha_i^l denoting the ll-th row of AiA_i, and εi\varepsilon_i is a group-level residual tolerance. The factor ∣Si∣p|S_i|^p normalizes the objective across groups of varying sizes.

    In practice, dictionary learning is performed using the convex ℓ1,2\ell_{1,2} norm ((p,q)=(1,2))((p,q) = (1,2)), which can be optimized online, while signal reconstruction is performed using the ℓ0,∞\ell_{0,\infty} pseudo-norm ((p,q)=(0,∞))((p,q) = (0,\infty)), counting the number of nonzero rows in AiA_i and solved approximately via simultaneous orthogonal matching pursuit.

  2. Knowl 2 — Patch Similarity Selection and Group Residual Tolerance

    equation

    In Learned Simultaneous Sparse Coding (LSSC), the set SiS_i of patches similar to the reference patch yi∈Rmy_i \in \mathbb{R}^m is constructed according to Euclidean distance:

    Si≜{j∈{1,…,n}∣∥yi−yj∥22≤ξ}S_i \triangleq \{j \in \{1, \dots, n\} \mid \|y_i - y_j\|_2^2 \le \xi\}

    where ξ\xi is a distance threshold set empirically for pixel intensities in the range [0,255][0, 255] by:

    ξ=(32σ)2m\xi = \frac{(32\sigma)^2}{m}

    with σ\sigma being the noise standard deviation and mm the patch dimension (number of pixels in a patch).

    For each group SiS_i, the residual fitting error threshold εi\varepsilon_i in the joint reconstruction constraint ∑j∈Si∥yj−Dαij∥22≤εi\sum_{j \in S_i} \|y_j - D\alpha_{ij}\|_2^2 \le \varepsilon_i is determined based on the statistical behavior of Gaussian noise. Because the normalized residual sum ∑j∈Si∥yj−Dαij∥22/σ2\sum_{j \in S_i} \|y_j - D\alpha_{ij}\|_2^2 / \sigma^2 follows a chi-squared distribution χm∣Si∣2\chi^2_{m|S_i|} with m∣Si∣m|S_i| degrees of freedom, the tolerance is computed as:

    εi=σ2Fm∣Si∣−1(τ)\varepsilon_i = \sigma^2 F_{m|S_i|}^{-1}(\tau)

    where Fm∣Si∣F_{m|S_i|} is the cumulative distribution function of χm∣Si∣2\chi^2_{m|S_i|}, Fm∣Si∣−1F_{m|S_i|}^{-1} is its inverse, and τ∈(0,1)\tau \in (0, 1) is a confidence threshold set to τ=0.8\tau = 0.8 for synthetic and real denoising experiments.

  3. Knowl 3 — Pixel Averaging and Image Reconstruction in Simultaneous Sparse Coding

    equation

    Once the dictionary D∈Rm×kD \in \mathbb{R}^{m \times k} and the sparse code vectors αij∈Rk\alpha_{ij} \in \mathbb{R}^k for all patch groups SiS_i have been computed, every pixel in the image receives multiple estimates coming from all patches and similarity groups containing that pixel. The final reconstructed image x∈Rnx \in \mathbb{R}^n is obtained by taking the weighted average of all reconstructed patch estimates:

    x=diag⁡(∑i=1n∑j∈SiRj1m)−1∑i=1n∑j∈SiRjDαijx = \operatorname{diag}\left(\sum_{i=1}^n \sum_{j \in S_i} R_j \mathbf{1}_m\right)^{-1} \sum_{i=1}^n \sum_{j \in S_i} R_j D \alpha_{ij}

    where:

    • Rj∈Rn×mR_j \in \mathbb{R}^{n \times m} is the binary placement matrix that maps a patch of size mm at position jj into its corresponding coordinates in the full image vector of size nn.
    • 1m∈Rm\mathbf{1}_m \in \mathbb{R}^m is an mm-dimensional vector filled with ones.
    • diag⁡(∑i=1n∑j∈SiRj1m)\operatorname{diag}\left(\sum_{i=1}^n \sum_{j \in S_i} R_j \mathbf{1}_m\right) is an n×nn \times n diagonal matrix whose ll-th entry equals the total number of estimates generated for pixel ll.
    • When each group is a singleton (Si={i}S_i = \{i\}), this formulation reduces to standard patch-averaging sparse coding.
  4. Knowl 4 — Simultaneous Sparse Coding Demosaicking Algorithm

    algorithm

    The demosaicking algorithm reconstructs a full-color RGB image from a single-channel Bayer-mosaicked sensor image yy by combining an offline natural image dictionary D0D_0 with an image-adapted dictionary D1D_1 under simultaneous sparse coding constraints.

    Input: Mosaicked image vector y∈Rny \in \mathbb{R}^n, Bayer sampling binary diagonal masks Mj∈Rm×mM_j \in \mathbb{R}^{m \times m} for each patch jj, offline pre-trained dictionary D0∈Rm×kD_0 \in \mathbb{R}^{m \times k}, similarity threshold ξ\xi, regularization parameters εi\varepsilon_i
    Output: Demosaicked full-color image x∈Rnx \in \mathbb{R}^n
    1. Group similar patches: For each patch yiy_i in yy, determine its similarity cluster Si={j∣∥yi−yj∥22≤ξ}S_i = \{j \mid \|y_i - y_j\|_2^2 \le \xi\}
    2. Initial reconstruction with D0D_0: For all i∈{1,…,n}i \in \{1, \dots, n\}, solve for Ai=[αij]j∈SiA_i = [\alpha_{ij}]_{j \in S_i}:
       min⁡Ai∈Rk×∣Si∣∥Ai∥0,∞s.t.∀j∈Si,Mj(yj−D0αij)=0\min_{A_i \in \mathbb{R}^{k \times |S_i|}} \|A_i\|_{0,\infty} \quad \text{s.t.} \quad \forall j \in S_i, M_j(y_j - D_0 \alpha_{ij}) = 0
       Average the resulting patch estimates via placement matrices RjR_j to obtain an intermediate demosaicked estimate xˉ\bar{x}
    3. Dictionary adaptation: On the estimate xˉ\bar{x}, train an adaptive dictionary D1∈Rm×kD_1 \in \mathbb{R}^{m \times k} using the LSSC formulation with large group error tolerances εi\varepsilon_i
    4. Final reconstruction with concatenated dictionary: Construct D2=[D0,D1]∈Rm×2kD_2 = [D_0, D_1] \in \mathbb{R}^{m \times 2k} and solve for each cluster:
       min⁡Ai∥Ai∥0,∞s.t.∀j∈Si,Mj(yj−D2αij)=0\min_{A_i} \|A_i\|_{0,\infty} \quad \text{s.t.} \quad \forall j \in S_i, M_j(y_j - D_2 \alpha_{ij}) = 0
    5. Aggregate estimates: Compute the final demosaicked image xx via weighted averaging:
       x=diag⁡(∑i=1n∑j∈SiRj1m)−1∑i=1n∑j∈SiRjD2αijx = \operatorname{diag}\left(\sum_{i=1}^n \sum_{j \in S_i} R_j \mathbf{1}_m\right)^{-1} \sum_{i=1}^n \sum_{j \in S_i} R_j D_2 \alpha_{ij}
    return xx

    Typical hyperparameters are patch size m=8×8m = 8 \times 8 (with 3 color channels, m=8×8×3=192m = 8 \times 8 \times 3 = 192), dictionary size k=256k = 256, and ξ=3×104\xi = 3 \times 10^4 for pixel values in [0,255][0, 255].

  5. Knowl 5 — Implementation Optimizations for Scalable Simultaneous Sparse Coding

    model/method

    Directly solving simultaneous sparse coding across all pairs of nn patches scales quadratically (n2n^2 code vectors αij\alpha_{ij}). The following algorithmic components reduce the computational complexity and memory footprint to scale linearly in nn:

    1. Semi-local grouping: Patch similarity search for yiy_i is restricted to a local window of size w×ww \times w around pixel ii (with w≤64w \le 64), reducing the worst-case number of code vectors to nw2n w^2.
    2. Disjoint patch clustering: Pixels are partitioned into disjoint clusters CkC_k such that all pixels i∈Cki \in C_k share an identical patch set SiS_i. Each pixel belongs to exactly one cluster, requiring the computation of only nn code vectors αij\alpha_{ij} in total.
    3. Offline dictionary pre-training: Initial dictionaries D0D_0 are pre-computed offline on 2×1072 \times 10^7 natural image patches randomly sampled from the 10,000 images of PASCAL VOC'07 using online dictionary learning, providing a superior starting point compared to smaller batch methods.
    4. Two-stage matching: An initial baseline denoising step (such as classical patch-wise sparse coding) is applied to the noisy image before patch grouping, significantly improving patch matching accuracy.
    5. Patch mean normalization: The mean intensity (or RGB color vector) of each patch is subtracted prior to sparse coding and re-added after estimation, improving numerical stability and visual quality.
    6. On-the-fly code evaluation: Online dictionary updates do not store codes globally in memory; the maximum number of code vectors stored at any time is bounded by the size of the largest cluster ∣Ck∣|C_k|.
  6. Knowl 6 — Denoising Benchmark Comparison on Additive White Gaussian Noise

    data/table

    Denoising performance was quantitatively benchmarked on 12 standard grayscale test images corrupted by synthetic zero-mean additive white Gaussian noise with standard deviation σ∈{5,10,15,20,25,50,100}\sigma \in \{5, 10, 15, 20, 25, 50, 100\}. Performance was evaluated using Peak Signal-to-Noise Ratio (PSNR in dB, defined as PSNR=10log⁡10(2552/MSE)\text{PSNR} = 10 \log_{10}(255^2 / \text{MSE})), averaged over 5 independent noise realizations.

    The evaluated methods include:

    • GSM: Gaussian Scale Mixtures in the wavelet domain (Portilla et al., 2003)
    • FoE: Fields of Experts (Roth & Black, 2005)
    • K-SVD: Adaptive sparse coding via K-SVD (Elad & Aharon, 2006)
    • BM3D: Block-Matching and 3D Filtering (Dabov et al., 2007)
    • SC: Fixed dictionary learned offline on 2×1072 \times 10^7 PASCAL VOC'07 patches, without grouping
    • LSC: Learned sparse coding adapted to the test image via ℓ1\ell_1 dictionary learning
    • LSSC: Learned simultaneous sparse coding combining learned dictionaries with joint patch sparsity
    σ\sigma GSM FoE K-SVD BM3D SC LSC LSSC
    5 37.05 37.03 37.42 37.62 37.46 37.66 37.67
    10 33.34 33.11 33.62 34.00 33.76 33.98 34.06
    15 31.31 30.99 31.58 32.05 31.72 31.99 32.12
    20 29.91 29.62 30.18 30.73 30.29 30.60 30.78
    25 28.84 28.36 29.10 29.72 29.18 29.52 29.74
    50 25.66 24.36 25.61 26.38 25.83 26.18 26.57
    100 22.80 21.36 22.10 23.25 22.46 22.62 23.39

    LSSC matches or outperforms BM3D across all noise levels, with the performance advantage increasing at higher noise levels (e.g., +0.19 dB over BM3D and +0.96 dB over K-SVD at σ=50\sigma = 50; +0.14 dB over BM3D and +1.29 dB over K-SVD at σ=100\sigma = 100).

  7. Knowl 7 — Per-Image Denoising Performance of LSSC on Standard Benchmarks

    data/table

    The quantitative denoising performance of Learned Simultaneous Sparse Coding (LSSC) was measured across 12 standard benchmark images at noise standard deviations σ∈{5,10,15,20,25,50,100}\sigma \in \{5, 10, 15, 20, 25, 50, 100\}. Hyperparameters used were k=512k = 512 atoms; patch sizes m=9×9m = 9 \times 9 for σ≤25\sigma \le 25, m=12×12m = 12 \times 12 for σ=50\sigma = 50, and m=16×16m = 16 \times 16 for σ=100\sigma = 100; τ=0.8\tau = 0.8; and ξ=(32σ)2/m\xi = (32\sigma)^2/m. Results are given in PSNR (dB), averaged over 5 noise realizations.

    Image σ=5\sigma = 5 σ=10\sigma = 10 σ=15\sigma = 15 σ=20\sigma = 20 σ=25\sigma = 25 σ=50\sigma = 50 σ=100\sigma = 100
    house 39.93 36.96 35.35 34.16 33.15 30.04 25.83
    peppers 38.18 34.80 32.82 31.37 30.21 26.62 23.00
    cameraman 38.32 34.21 32.01 30.57 29.51 26.42 23.08
    lena 38.69 35.83 34.15 32.90 31.87 28.87 25.82
    barbara 38.48 34.97 33.00 31.57 30.47 27.06 23.59
    boat 37.35 34.02 32.20 30.89 29.87 26.74 23.84
    hill 37.17 33.67 31.89 30.71 29.80 27.05 24.44
    couple 37.45 33.98 32.06 30.69 29.61 26.30 23.28
    man 37.89 34.06 32.01 30.64 29.63 26.69 24.00
    fingerprint 36.70 32.57 30.31 28.78 27.62 24.25 21.26
    bridge 35.78 31.22 28.92 27.46 26.42 23.68 21.46
    flintstones 36.13 32.46 30.78 29.63 28.71 25.16 21.10
    Average 37.67 34.06 32.12 30.78 29.74 26.57 23.39
  8. Knowl 8 — Color Demosaicking Benchmark Performance on Kodak PhotoCD Dataset

    data/table

    Color demosaicking performance was evaluated on the 24 RGB images (512×768512 \times 768 pixels) of the standard Kodak PhotoCD benchmark with a Bayer color filter mask applied. Following standard benchmarking protocols (Paliy et al., 2007), a 15-pixel border was excluded from the evaluation to prevent boundary bias. Results are reported in PSNR (dB).

    The compared algorithms are:

    • AP: Alternating Projections (Gunturk et al., 2002)
    • DL: Directional Linear Minimum Mean Square-Error Estimation (Zhang & Wu, 2005)
    • LPA: Local Polynomial Approximation (Paliy et al., 2007)
    • SC: Sparse Coding with fixed offline dictionary (Mairal et al., 2009)
    • LSC: Learned Sparse Coding adapted to the image (Mairal et al., 2009)
    • LSSC: Learned Simultaneous Sparse Coding (Mairal et al., 2009)
    Image AP DL LPA SC LSC LSSC
    1 37.84 38.46 40.47 40.84 40.92 41.36
    2 39.64 40.89 41.36 41.76 42.03 42.24
    3 41.40 42.66 43.47 43.15 43.92 44.24
    4 39.92 40.49 40.84 41.99 42.14 42.45
    5 37.28 38.07 37.51 38.72 39.15 39.45
    6 38.69 40.19 40.92 41.29 41.36 41.71
    7 41.75 42.35 43.06 43.30 43.59 44.06
    8 35.58 36.02 37.13 37.42 37.38 37.57
    9 41.84 43.05 43.50 43.17 43.74 43.83
    10 41.93 42.54 42.77 43.01 43.17 43.33
    11 39.25 40.01 40.51 41.19 41.29 41.51
    12 42.62 43.45 44.01 44.29 44.49 44.90
    13 34.28 34.75 36.08 36.16 36.29 36.35
    14 35.66 36.91 36.86 37.64 38.48 38.77
    15 39.17 39.82 40.09 41.04 41.24 41.74
    16 42.10 43.75 44.02 44.36 44.42 44.91
    17 41.23 41.68 41.75 41.75 41.86 41.98
    18 37.31 37.64 37.59 38.05 38.27 38.38
    19 39.99 41.01 41.55 41.58 41.71 42.31
    20 40.63 41.24 41.48 41.95 42.25 42.27
    21 38.72 39.10 39.61 40.55 40.59 40.65
    22 37.63 38.37 38.44 38.73 38.97 39.24
    23 41.93 43.22 43.92 43.47 43.93 44.34
    24 34.74 35.55 35.44 35.59 35.85 35.89
    Average 39.21 40.05 40.52 40.88 41.13 41.39

    LSSC outperforms the state-of-the-art LPA method by 0.87 dB on average. When evaluated over the full image (including borders), SC achieves 40.72 dB, LSC achieves 40.98 dB, and LSSC achieves 41.24 dB (compared to 39.56 dB for fixed-dictionary SC and 40.32 dB for adaptive LSC in Mairal et al., 2008), while eliminating typical demosaicking color fringe and zipper artefacts.

  9. Knowl 9 — Raw Sensor Denoising and Processing Pipeline for Digital Cameras

    model/method

    To restore images corrupted by real sensor noise (e.g., Canon Powershot G9 at 1600 ISO), denoising is performed directly on the raw CFA (Color Filter Array) mosaicked sensor data before running internal camera rendering pipelines.

    The pipeline operates in five successive steps:

    1. Raw data extraction: Extract the raw un-demosaicked Bayer array directly from the camera sensor data (e.g., using dcraw).
    2. Channel scaling: Manually or calibrationally scale the R, G, and B sensor values so that their noise variances visually match, producing approximately uniform noise across the mosaicked array.
    3. Mosaicked denoising: Apply LSSC denoising directly to the single-channel mosaicked Bayer array using patch size m=8×8m = 8 \times 8, k=256k = 256 atoms, user-estimated uniform noise standard deviation σ\sigma, and threshold ξ=(32σ)2/m\xi = (32\sigma)^2/m.
    4. Demosaicking: Apply the simultaneous sparse coding demosaicking algorithm to reconstruct full RGB triplets at every pixel.
    5. Color post-processing: Apply white balancing, color space conversion to sRGB, gamma correction, and contrast enhancement.

    Denoising the raw mosaicked array prior to demosaicking avoids amplifying sensor noise through interpolation and non-linear color transformations, yielding sharper results with fewer artefacts than commercial pipelines (e.g., NoiseWare 4.2 or Adobe Camera Raw).

  10. Knowl 10 — Spatially Homogeneous Noise Assumption in Raw Camera Restoration

    limitation

    The restoration model assumes that noise in the raw sensor data is independent and identically distributed additive white Gaussian noise with a single uniform standard deviation σ\sigma across the entire sensor grid. In physical digital cameras, however, sensor noise is signal-dependent, spatially non-uniform, and channel-dependent (varying with pixel brightness, ISO amplification, and sensor temperature).

    Although scaling the R, G, and B channels mitigates global variance differences, assuming strict spatial uniformity can lead to suboptimal restoration in certain image regions (for instance, partially reconstructing noise in flat background areas of high-ISO photographs).

Coverage note — None was omitted; all primary contributions (LSSC formulation, grouping and tolerance equations, pixel reconstruction, demosaicking algorithm, implementation speedups, synthetic denoising benchmarks, Kodak demosaicking benchmarks, real RAW camera restoration pipeline, and noise model limitations) are included.

References

  1. 1.S. Awate and R. Whitaker. Unsupervised, information-theoretic, adaptive image filtering for image restoration. IEEE T. PAMI, 364–376, 2006.
  2. 2.M. Bertalmio, G. Sapiro, V. Caselles, and C. Ballester. Image inpainting. In Proc. Comp. Graph. and Interact. Tech., 2000.
  3. 3.A. Buades, B. Coll, and J. Morel. A non-local algorithm for image denoising. In Proc. IEEE CVPR, 2005.
  4. 4.A. Buades, B. Coll, J. Morel, and C. Sbert. Non-local demosaicing. Technical report, 2007. Preprint CMLA 2007-15.
  5. 5.S. Chen, D. Donoho, and M. Saunders. Atomic decomposition by basis pursuit. SIAM J. Sc. Comp., 20:33–61, 1999.
  6. 6.A. Criminisi, P. Pérez, and K. Toyama. Region filling and object removal by exemplar-based inpainting. IEEE T. IP, 13(9):1200–1212, 2004.
  7. 7.K. Dabov, A. Foi, V. Katkovnik, and K. Egiazarian. Image Denoising by Sparse 3-D Transform-Domain Collaborative Filtering. IEEE T. IP, 16(8):2080–2095, 2007.
  8. 8.D.L. Donoho, I.M. Johnstone. Adapting to Unknown Smoothness Via Wavelet Shrinkage. J. of the American Stat. Assoc., 90(432):1200–1224, 1995.
  9. 9.B. Efron, T. Hastie, I. Johnstone, and R. Tibshirani. Least angle regression. Ann. Statist., 32(2):407–499, 2004.
  10. 10.A. Efros and T. Leung. Texture synthesis by non-parametric sampling. In Proc. ICCV, 1999.
  11. 11.M. Elad and M. Aharon. Image denoising via sparse and redundant representations over learned dictionaries. IEEE T. IP, 54(12):3736–3745, 2006.
  12. 12.J. Friedman, T. Hastie, H. Hölfling, and R. Tibshirani. Pathwise coordinate optimization. Ann. Stats., 1(2):302–332, 2007.
  13. 13.B. Gunturk, Y. Altunbasak, and R. Mersereau. Color plane interpolation using alternating projections. IEEE T. IP, 11(9):997–1013, 2002.
  14. 14.Y. Li and D. P. Huttenlocher. Sparse long-range random field and its application to image denoising. In Proc. ECCV, 2008.
  15. 15.J. Mairal, M. Elad, and G. Sapiro. Sparse representation for color image restoration. IEEE T. IP, 17(1):53–69, 2008.
  16. 16.J. Mairal, F. Bach, J. Ponce, and G. Sapiro. Online dictionary learning for sparse coding. ICML, 2009.
  17. 17.S. Mallat. A Wavelet Tour of Signal Processing, 2nd Edition. Academic Press, 1999.
  18. 18.S. Mallat and Z. Zhang. Matching pursuit in a time-frequency dictionary. IEEE T. SP, 41(12):3397–3415, 1993.
  19. 19.B. A. Olshausen and D. J. Field. Sparse coding with an overcomplete basis set: A strategy employed by V1? Vision Research, 37:3311–3325, 1997.
  20. 20.D. Paliy, V. Katkovnik, R. Bilcu, S. Alenius, and K. Egiazarian. Spatially adaptive color filter array interpolation for noiseless and noisy data. Int. J. Imag. Sys. Tech., 17(3), 2007.
  21. 21.P. Perona and J. Malik. Scale-space and edge detection using anisotropic diffusion. IEEE T. PAMI, 12(7):629–639, 1990.
  22. 22.J. Portilla, V. Strela, M. Wainwright, and E. Simoncelli. Image denoising using scale mixtures of Gaussians in the wavelet domain. IEEE T. IP, 12(11):1338–1351, 2003.
  23. 23.R. Puetter, T. Gosnell, and A. Yahil. Digital image reconstruction: deblurring and denoising. Annu. Rev. Astron. Astrophys., 43, 2005.
  24. 24.S. Roth and M. J. Black. Fields of experts: A framework for learning image priors. In Proc. IEEE CVPR, 2005.
  25. 25.L. Rudin and S. Osher. Total variation based image restoration with free local constraints. In Proc. ICIP, 1994.
  26. 26.A. Szlam, M. Maggioni, and R. Coifman. Regularization on graphs with function-adapted diffusion processes. JMLR, 2007.
  27. 27.R. Tibshirani. Regression shrinkage and selection via the lasso. J. Royal. Statist. Soc. B, 58(1):267–288, 1996.
  28. 28.J. A. Tropp. Algorithms for simultaneous sparse approximation. Sig. Proc., 86:572–602, 2006.
  29. 29.B. Turlach, W. Venables, and S. Wright. Simultaneous variable selection. Technometrics, 47(3):349, 2005.
  30. 30.S. Weisberg. Applied Linear Regression. Wiley, 1980.
  31. 31.M. Yuan and Y. Lin. Model selection and estimation in regression with grouped variables. J. Royal. Statist. Soc. B, 68(1):49–67, 2006.
  32. 32.L. Zhang and X. Wu. Color demosaicking via directional linear minimum mean square-error estimation. IEEE T. IP, 14(12):2167–2178, 2005.

Citation

MLA
Jenatton, R., et al. “Structured Variable Selection with Sparsity-Inducing Norms”. Journal of Machine Learning Research 12 (2011) 2777-2824, 2009, http://arxiv.org/abs/0904.3523v3.
APA
Jenatton, R., Audibert, J.-Y., & Bach, F. (2009). Structured Variable Selection with Sparsity-Inducing Norms. Journal of Machine Learning Research 12 (2011) 2777-2824. http://arxiv.org/abs/0904.3523v3
Chicago
Jenatton, R., J.-Y. Audibert, and F. Bach. 2009. “Structured Variable Selection with Sparsity-Inducing Norms”. Journal of Machine Learning Research 12 (2011) 2777-2824. http://arxiv.org/abs/0904.3523v3.
Harvard
Jenatton, R., Audibert, J.-Y. and Bach, F. (2009) “Structured Variable Selection with Sparsity-Inducing Norms”, Journal of Machine Learning Research 12 (2011) 2777-2824 [Preprint]. Available at: http://arxiv.org/abs/0904.3523v3.
Vancouver
1. Jenatton R, Audibert J-Y, Bach F (2009) Structured Variable Selection with Sparsity-Inducing Norms. Journal of Machine Learning Research 12 (2011) 2777-2824

BibTeX

@article{jenatton2009structured,
  title = {Structured Variable Selection with Sparsity-Inducing Norms},
  author = {Jenatton, Rodolphe and Audibert, Jean-Yves and Bach, Francis},
  year = {2009},
  journal = {Journal of Machine Learning Research 12 (2011) 2777-2824},
  url = {http://arxiv.org/abs/0904.3523v3},
  eprint = {0904.3523}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF