Super-resolution from a single image

Daniel GlasnerShai BagonMichal Irani

article2009ICCV2,028 citations

Unifies classical multi-image reconstruction and example-based learning into a single framework that generates high-resolution details from a lone input image by exploiting cross-scale and within-scale patch redundancies.

Listen

Recovering high-resolution detail from low-resolution imagery is a persistent challenge across visual computing, defense, intelligence, and consumer media. Existing techniques historically faced major trade-offs: classical multi-image super-resolution relies on multiple precisely aligned low-resolution inputs and fails beyond modest magnification limits (typically less than a factor of 2), while example-based methods require massive external image training databases that can introduce artificial, erroneous features ("hallucinations"). The article aims to evaluate and demonstrate a unified super-resolution framework capable of generating high-quality image enhancements from as little as a single low-resolution input, without relying on external training databases or prior examples.

The authors develop an approach centered on the natural property of patch redundancy, showing that small image fragments (e.g., 5×5 pixels) naturally repeat both within the same image scale and across different down-scaled representations of the image. The article validates this premise through statistical analysis on 300 natural images from the Berkeley Segmentation Database. By treating recurring within-scale patches as independent observations, the framework establishes classical reconstruction constraints; concurrently, cross-scale patch matches provide natural high-to-low resolution exemplar pairs. The system solves the combined system of equations across a coarse-to-fine scale cascade, using iterative back-projection to ensure strict consistency with the original low-resolution image.

The findings confirm strong empirical and practical performance across several dimensions. First, the statistical analysis reveals high internal redundancy: over 90% of patches in natural images recur at least nine times within their native scale, and over 80% of edge and texture patches recur across significant scale reductions (down to 41% of original size). Second, combining within-scale and cross-scale constraints successfully reconstructs genuine high-frequency details (such as fine railings and distinct text) that standard interpolation and classical single-image methods blur or miss. Third, while the cross-scale example-based component provides the primary resolution boost, the classical constraint component effectively anchors the solution to the input data, preventing erroneous feature hallucinations. Finally, qualitative and quantitative benchmarks demonstrate that this self-contained single-image method achieves visual quality comparable to or exceeding leading external database-driven and edge-modeled approaches at magnification factors up to 4x.

These findings indicate that organizations can achieve state-of-the-art super-resolution without collecting, curating, or maintaining massive external reference datasets, thereby lowering data-storage costs and eliminating privacy or licensing risks associated with third-party training data. Furthermore, because the framework adaptively adjusts based on localized patch repetition, it reduces the risk of generating false artifacts in safety- or compliance-critical imaging domains.

Stakeholders seeking to deploy super-resolution workflows should consider adopting internal self-similarity pipelines as a lightweight alternative to database-heavy models, particularly when operating on standalone images. The source also supports extending this unified formulation into existing multi-frame video pipelines or hybrid database architectures to maximize reconstruction fidelity. Practitioners should note that the primary limitation occurs in image regions lacking cross-scale self-similarity—such as unique, non-repeating margin text—where the method gracefully falls back to classical enhancement limits without inventing artificial details. Overall confidence in the underlying methodology is high for natural imagery featuring standard textural and perspective redundancies.

Cover for Super-resolution from a single image

Abstract

Methods for super-resolution can be broadly classified into two families of methods: (i) The classical multi-image super-resolution (combining images obtained at subpixel misalignments), and (ii) Example-Based super-resolution (learning correspondence between low and high resolution image patches from a database). In this paper we propose a unified framework for combining these two families of methods. We further show how this combined approach can be applied to obtain super resolution from as little as a single image (with no database or prior examples). Our approach is based on the observation that patches in a natural image tend to redundantly recur many times inside the image, both within the same scale, as well as across different scales. Recurrence of patches within the same image scale (at subpixel misalignments) gives rise to the classical super-resolution, whereas recurrence of patches across different scales of the same image gives rise to example-based super-resolution. Our approach attempts to recover at each pixel its best possible resolution increase based on its patch redundancy within and across scales.

Table of Contents

  • 1. Introduction
  • 2. Patch Redundancy in a Single Image
  • 3. Single Image SR - A Unified Framework
  • 3.1. Employing in-scale patch redundancy
  • 3.2. Employing cross-scale patch redundancy
  • 3.3. Combining Classical and Example-Based SR
  • 4. Experimental Results
  • References

Knowls

  1. Knowl 1 — Unified Single-Image Super-Resolution Formulation

    model/method

    The unified super-resolution (SR) framework simultaneously incorporates classical multi-image reconstruction constraints and database-free example-based constraints into a single linear system.

    Let L=I0L = I_0 be the input low-resolution image and H=InH = I_n be the unknown target high-resolution image, related by the point spread function (PSF) blur kernel BB and total scale factor ss: L=(H∗B)↓sL = (H * B) \downarrow_s, where ↓\downarrow denotes subsampling. A continuous scale progression is discretized into a cascade of images I0,I1,…,In=HI_0, I_1, \dots, I_n = H with scale factors sl=αls_l = \alpha^l relative to LL for a fixed step α>1\alpha > 1, and associated blur kernels BlB_l.

    Each patch QlQ_l extracted at resolution level l∈{0,1,…,n−1}l \in \{0, 1, \dots, n-1\} (either within-scale for l=0l=0 or transferred from cross-scale parent patches for l>0l > 0) induces a set of linear equations on the target high-resolution pixels H(q)H(q):

    Ql(p)=(H∗Bn−l)↓sn−l(p)=∑qi∈Support(Bn−l)H(qi)Bn−l(qi−p)Q_l(p) = (H * B_{n-l}) \downarrow_{s_{n-l}}(p) = \sum_{q_i \in \text{Support}(B_{n-l})} H(q_i) B_{n-l}(q_i - p)

    where Bn−lB_{n-l} is a compactly supported blur kernel accounting for the residual scale gap (n−l)(n-l) between intermediate level ll and the target level nn. Each equation is weighted by a reliability score based on the patch similarity. As l→nl \to n, Bn−lB_{n-l} approaches the Dirac delta function δ\delta, improving the conditioning of the linear system. If only within-scale matches are found (l=0l=0), the system reduces to classical single-image SR; if no matches are found, it reduces to localized deblurring.

  2. Knowl 2 — Cross-Scale Patch Recurrence for Internal Example-Based Super-Resolution

    model/method

    Example-based super-resolution (SR) learns low-to-high resolution patch mappings without an external training database by exploiting patch recurrence across different scales within the same image.

    A downsampled pyramid of known images is constructed from the low-resolution input image L=I0L = I_0: I−l=(L∗Bl)↓slI_{-l} = (L * B_l) \downarrow_{s_l} for scale levels l=1,…,ml = 1, \dots, m, where BlB_l is the downscaling blur kernel and sl=αls_l = \alpha^l.

    For each pixel pp in LL and its surrounding patch P0(p)P_0(p) (typically 5×55 \times 5 pixels):

    1. An Approximate Nearest Neighbor search locates the most similar patch P−l(p~)P_{-l}(\tilde{p}) in a coarser image I−lI_{-l}.
    2. The higher-resolution 'parent' patch Q0(sl⋅p~)Q_0(s_l \cdot \tilde{p}) corresponding to P−l(p~)P_{-l}(\tilde{p}) is extracted directly from the input image LL.
    3. The pair [P−l(p~),Q0(sl⋅p~)][P_{-l}(\tilde{p}), Q_0(s_l \cdot \tilde{p})] forms an internal low-res/high-res correspondence, providing an estimate Ql(sl⋅p)Q_l(s_l \cdot p) for the unknown high-resolution parent patch of P0(p)P_0(p) at intermediate scale level ll in the target cascade.
  3. Knowl 3 — Coarse-to-Fine Single-Image Super-Resolution with Iterative Back-Projection

    algorithm

    The unified single-image super-resolution algorithm reconstructs the high-resolution target image HH incrementally across scale levels l=0,…,n−1l = 0, \dots, n-1 using internal patch correspondences and back-projection consistency.

    Input: Low-resolution image LL, scaling factor step α\alpha (e.g., α=1.25\alpha = 1.25 or α=21/3\alpha = 2^{1/3}), target level nn, patch size (e.g., 5×55 \times 5), number of neighbors k=9k = 9.
    Output: High-resolution image H=InH = I_n.
    Construct down-scaled image pyramid I−m,…,I−1I_{-m}, \dots, I_{-1} from LL via I−l=(L∗Bl)↓slI_{-l} = (L * B_l) \downarrow_{s_l}.
    Initialize I0=LI_0 = L.
    for l=0l = 0 to n−1n-1 do
        For each patch in IlI_l, search for kk nearest neighbors across known resolution levels {I−m,…,Il}\{I_{-m}, \dots, I_l\} using Approximate Nearest Neighbors.
        Extract parent patches from the appropriate higher-resolution levels to form candidate high-resolution patches QQ.
        Formulate the linear system of equations for Il+1I_{l+1} where each patch QQ induces constraints Q=(Il+1∗Bres)↓sresQ = (I_{l+1} * B_{\text{res}}) \downarrow_{s_{\text{res}}}, weighted by patch similarity scores.
        Solve the weighted linear system to recover an initial estimate of Il+1I_{l+1}.
        Simulate low-resolution image: Lsim=(Il+1∗Bl+1)↓sl+1L_{\text{sim}} = (I_{l+1} * B_{l+1}) \downarrow_{s_{l+1}}.
        Compute error residual: E=L−LsimE = L - L_{\text{sim}}.
        Back-project residual EE onto Il+1I_{l+1} to enforce consistency with the input LL.
    end for
    return H=InH = I_n
  4. Knowl 4 — Empirical Patch Recurrence Statistics Across Image Scales

    empirical result

    An empirical evaluation on 300 natural grayscale images from the Berkeley Segmentation Database measures the frequency of small (5×55 \times 5) patch repetitions within the same scale and across scaled-down resolutions (IsI_s scaled by 1.25s1.25^s for s=0,−1,…,−6s = 0, -1, \dots, -6, down to 0.26=1.25−60.26 = 1.25^{-6} of the original dimensions).

    When evaluating all patches across the database:

    • Over 90%90\% of patches have ≥9\ge 9 similar patches within the original scale (s=0s = 0).
    • Over 80%80\% of patches have ≥9\ge 9 similar patches at scale factor 0.41=1.25−40.41 = 1.25^{-4}.
    • 70%70\% of patches have ≥9\ge 9 similar patches at scale factor 0.26=1.25−60.26 = 1.25^{-6}.

    When evaluating only the top 25%25\% highest-variance patches (retaining edges, corners, and texture while removing uniform regions):

    • Over 80%80\% of high-variance patches have ≥9\ge 9 similar patches within the original scale.
    • Over 70%70\% have ≥9\ge 9 similar patches at scale factor 0.410.41.
    • 60%60\% have ≥9\ge 9 similar patches at scale factor 0.260.26.

    The lowest scale in which a patch recurs establishes the upper bound for its potential resolution increase using internal self-similarity alone.

  5. Knowl 5 — Patch Similarity Metric via Sub-Pixel Shift Self-Distance

    definition

    Patch similarity between two 5×55 \times 5 patches P1P_1 and P2P_2 is computed by measuring their Gaussian-weighted Sum of Squared Differences (SSD) after subtracting their respective mean intensity (DC component).

    Because textured and high-variance patches produce larger baseline SSD errors than smooth patches under minor misalignments, a patch-adaptive similarity threshold is defined for each source patch PP. The threshold distance dthresh(P)d_{\text{thresh}}(P) is the Gaussian-weighted SSD between PP and a copy of PP shifted by 0.50.5 pixels. A candidate patch P′P' is classified as similar to PP if and only if:

    d(P,P′)≤dthresh(P)=d(P,Pshift(0.5))d(P, P') \le d_{\text{thresh}}(P) = d\left(P, P_{\text{shift}(0.5)}\right)

  6. Knowl 6 — In-Scale Patch Redundancy as Classical Super-Resolution Constraints

    model/method

    Classical multi-image super-resolution relies on multiple low-resolution frames of the same scene captured with subpixel shifts to construct an overdetermined linear system for the unknown high-resolution pixels.

    In a single image L=(H∗B)↓sL = (H * B) \downarrow_s, recurring patches P1,…,PkP_1, \dots, P_k similar to a patch PP at pixel pp occur at subpixel misalignments (determined at 1/s1/s pixel accuracy, where ss is the magnification scale factor). These patches are treated as independent observations of the same underlying high-resolution scene structure. Each matching patch PjP_j induces a linear constraint on the high-resolution pixels H(q)H(q) within the support of blur kernel BB:

    Pj(p)=(H∗B)(q)=∑qi∈Support(B)H(qi)B(qi−q)P_j(p) = (H * B)(q) = \sum_{q_i \in \text{Support}(B)} H(q_i) B(q_i - q)

    Each constraint is weighted by the similarity score of PjP_j to PP. Combining these constraints for kk neighbors (typically k=9k=9) converts the underdetermined single-image problem into a determined linear system, achieving magnification up to factors of ≈1.6−2.0\approx 1.6 - 2.0 using only within-scale redundancy.

  7. Knowl 7 — Complementary Roles of Classical and Example-Based Constraints in Unified Super-Resolution

    theoretical result

    The unified super-resolution framework represents an optimization problem combining a data fidelity term and two distinct regularization priors:

    1. Data Term: The physical image degradation model (H∗B)↓s=L(H * B) \downarrow_s = L, requiring the reconstructed high-resolution image HH to remain consistent with the input low-resolution image LL.
    2. Example-Based Prior: Generated by cross-scale patch correspondences, this prior provides high-frequency details and edge sharpens beyond the Nyquist frequency of LL, overcoming the theoretical resolution limits (<2×<2\times) of classical reconstruction methods.
    3. Classical Multi-Patch Prior: Generated by within-scale subpixel patch recurrences, this prior enforces subpixel geometric consistency.

    The classical reconstruction constraints constrain the solution space and prevent the example-based component from hallucinating erroneous high-resolution textures or artifacts that are inconsistent with the observed low-resolution data.

  8. Knowl 8 — Color Channel Processing in Single-Image Super-Resolution

    model/method

    For super-resolution of color images, the input image is transformed from RGB color space to the YIQ color space.

    The unified single-image super-resolution algorithm is applied solely to the luminance (YY) channel, which contains the high-frequency structural content (edges, corners, and fine textures). The chromatic channels (II and QQ), which are dominated by low-frequency color variations, are magnified using standard bicubic interpolation. The super-resolved YY channel and the bicubic-interpolated II and QQ channels are then combined and converted back to RGB.

Coverage note — None was omitted; all key theoretical formulations, empirical redundancy measurements, patch matching thresholds, unified equations, algorithm steps, and color pipeline have been extracted.

References

  1. 1.S. Arya and D. M. Mount. Approximate nearest neighbor queries in fixed dimensions. In SODA, 1993.
  2. 2.S. Baker and T. Kanade. Hallucinating faces. In Automatic Face and Gesture Recognition, 2000.
  3. 3.S. Baker and T. Kanade. Limits on super-resolution and how to break them. PAMI, (9), 2002.
  4. 4.A. Buades, B. Coll, and J. M. Morel. A review of image denoising algorithms, with a new one. SIAM MMS, (2), 2005.
  5. 5.D. Capel. Image Mosaicing and Super-Resolution. Springer–Verlag, 2004.
  6. 6.M. Ebrahimi and E. Vrscay. Solving the inverse problem of image zooming using "self-examples". In Image Analysis and Recognition, 2007.
  7. 7.A. A. Efros and T. K. Leung. Texture synthesis by non-parametric sampling. In ICCV, 1999.
  8. 8.S. Farsiu, M. Robinson, M. Elad, and P. Milanfar. Fast and robust multiframe super resolution. T-IP, (10), 2004.
  9. 9.R. Fattal. Image upsampling via imposed edge statistics. In SIGGRAPH, 2007.
  10. 10.W. Freeman, E. Pasztor, and O. Carmichael. Learning low-level vision. IJCV, (1), 2000.
  11. 11.W. T. Freeman, T. R. Jones, and E. C. Pasztor. Example-based super-resolution. Comp. Graph. Appl., (2), 2002.
  12. 12.M. Irani and S. Peleg. Improving resolution by image registration. CVGIP, (3), 1991.
  13. 13.K. Kim and Y. Kwon. Example-based learning for single-image SR and JPEG artifact removal. MPI-TR, (173), 08.
  14. 14.Z. Lin and H. Shum. Fundamental Limits of Reconstruction-Based Superresolution Algorithms under Local Translation. PAMI, (1), 04.
  15. 15.G. Peyre, S. Bougleux, and L. D. Cohen. Non-local regularization of inverse problems. In ECCV, 2008.
  16. 16.M. Protter, M. Elad, H. Takeda, and P. Milanfar. Generalizing the nonlocal-means to super-resolution reconstruction. T-IP, 09.
  17. 17.D. Simakov, Y. Caspi, E. Shechtman, and M. Irani. Summarizing visual data using bidirectional similarity. CVPR, 2008.
  18. 18.N. Suetake, M. Sakano, and E. Uchino. Image super-resolution based on local self-similarity. Opt.Rev., (1), 2008.
  19. 19.J. Sun, Z. Xu, and H. Shum. Image super-resolution using gradient profile prior. In CVPR, 2008.

Citation

MLA
Glasner, D., et al. “Super-resolution from a Single Image”. 2009 IEEE 12th International Conference on Computer Vision, 2009, pp. 349–56, https://doi.org/10.1109/ICCV.2009.5459271.
APA
Glasner, D., Bagon, S., & Irani, M. (2009). Super-resolution from a single image. 2009 IEEE 12th International Conference on Computer Vision, 349–356. https://doi.org/10.1109/ICCV.2009.5459271
Chicago
Glasner, D., S. Bagon, and M. Irani. 2009. “Super-resolution from a Single Image”. 2009 IEEE 12th International Conference on Computer Vision, 349–56. https://doi.org/10.1109/ICCV.2009.5459271.
Harvard
Glasner, D., Bagon, S. and Irani, M. (2009) “Super-resolution from a single image”, 2009 IEEE 12th International Conference on Computer Vision. IEEE, pp. 349–356. Available at: https://doi.org/10.1109/ICCV.2009.5459271.
Vancouver
1. Glasner D, Bagon S, Irani M (2009) Super-resolution from a single image. In: 2009 IEEE 12th International Conference on Computer Vision. IEEE, pp 349–356

BibTeX

@inproceedings{Glasner_2009, title={Super-resolution from a single image}, url={http://dx.doi.org/10.1109/ICCV.2009.5459271}, DOI={10.1109/iccv.2009.5459271}, booktitle={2009 IEEE 12th International Conference on Computer Vision}, publisher={IEEE}, author={Glasner, Daniel and Bagon, Shai and Irani, Michal}, year={2009}, month=Sept, pages={349–356} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF

License: IEEE