CoSeR: Bridging Image and Language for Cognitive Super-Resolution

Haoze SunWenbo LiJianzhuang LiuHaoyu ChenRenjing PeiXueyi ZouYouliang YanYujiu Yang

article2024CVPR88 citations

Proposes a cognitive super-resolution framework that combines image appearance with language comprehension to generate semantic reference images and guide diffusion models using an unified attention module for accurate, photorealistic detail restoration.

Listen

Real-world image super-resolution—the process of enhancing degraded, low-resolution images into sharp, high-resolution counterparts—is critical for downstream applications such as mobile photography, autonomous driving, and robotics. Conventional restoration techniques primarily rely on bottom-up, pixel-level processing that struggles to interpret the broader context of an image. As a result, existing systems often fail to restore severely degraded elements or introduce semantically incorrect textures, limiting their effectiveness in real-world scenarios.

The article demonstrates the Cognitive Super-Resolution framework, which integrates image appearance and language understanding to endow super-resolution models with holistic scene comprehension. The primary objective is to evaluate whether combining implicit diffusion priors with automatically generated reference images can significantly improve semantic accuracy and visual realism while maintaining fidelity to the original input.

To achieve this, the authors designed a cognitive encoder that extracts multi-token cognitive embeddings directly from low-resolution inputs without requiring large external captioning models. These embeddings activate the generative priors of a pre-trained text-to-image diffusion model and automatically synthesize semantically aligned, high-quality reference images. A unified attention architecture, termed All-in-Attention, then integrates the low-resolution input, the generated reference image, and the cognitive embedding directly into the diffusion process. The approach was evaluated on synthetic benchmarks using ImageNet as well as established real-world datasets, including RealSR and DRealSR, using standard perceptual and non-reference quality metrics alongside a human user study.

The experimental findings show that the proposed framework consistently outperforms existing state-of-the-art super-resolution methods across multiple benchmarks. In perceptual quality assessments, the model achieved a 13.8% improvement in Fréchet Inception Distance over the second-best approach on the ImageNet test set, alongside gains of 3.8% on RealSR and 4.7% on DRealSR. The multi-token cognitive adapter reduced semantic and textural bias compared to standard single-token approaches, producing reference images that closely mirror the ground truth. Furthermore, in an independent user study involving real-world low-resolution images, approximately 80% of participants selected the framework's outputs as visually superior to those of competing models.

These results demonstrate that mimicking human top-down cognition—understanding global context before refining fine details—substantially improves image restoration quality. The integration of automated reference generation eliminates the operational burden of manually sourcing high-definition reference images. Additionally, the unified attention mechanism provides a practical method to balance generative realism with strict input fidelity, reducing the risk of generating synthetic artifacts in high-stakes visual tasks.

Organizations evaluating advanced computer vision systems should consider adopting cognitive, diffusion-based restoration architectures for tasks requiring high visual fidelity. Development teams should incorporate lightweight cognitive adapters rather than bulky language-captioning pipelines to maintain parameter efficiency. Future efforts should focus on deploying the framework across diverse real-world edge devices, expanding evaluation to non-natural image domains, and streamlining inference sampling speeds to support time-critical applications.

While the model demonstrates high performance, potential users should note that the evaluation relied on synthetic degradations for training and fixed downscaling ratios during benchmark testing. A moderate level of caution is advised when processing imagery that substantially deviates from common photographic domains, though the strong quantitative results and user study provide high confidence in the framework's broad real-world applicability.

arXiv: 2311.16512
Cover for CoSeR: Bridging Image and Language for Cognitive Super-Resolution

Abstract

Existing super-resolution (SR) models primarily focus on restoring local texture details, often neglecting the global semantic information within the scene. This oversight can lead to the omission of crucial semantic details or the introduction of inaccurate textures during the recovery process. In our work, we introduce the Cognitive Super-Resolution (CoSeR) framework, empowering SR models with the capacity to comprehend low-resolution images. We achieve this by marrying image appearance and language understanding to generate a cognitive embedding, which not only activates prior information from large text-to-image diffusion models but also facilitates the generation of high-quality reference images to optimize the SR process. To further improve image fidelity, we propose a novel condition injection scheme called "All-in-Attention", consolidating all conditional information into a single module. Consequently, our method successfully restores semantically correct and photorealistic details, demonstrating state-of-the-art performance across multiple benchmarks. Project page: https://coser-main.github.io/

Table of Contents

  • 1. Introduction
  • 2. Related work
  • 2.1. Real-World Image Super-Resolution
  • 2.2. Diffusion-Based Super-Resolution
  • 2.3. Reference-Based Super-Resolution
  • 3. Methodology
  • 3.1. Cognitive Encoder
  • 3.2. Reference Image Generation and Encoding
  • 3.3. All-in-Attention Module
  • 4. Experiments
  • 4.1. Implementation Details
  • 4.2. Experimental Settings
  • 4.3. Comparison with State of the Arts
  • 4.4. Ablation Study
  • 5. Conclusion
  • References

Knowls

  1. Knowl 1 — CoSeR cognitive super-resolution pipeline

    model/method

    CoSeR is a two-stage real-world image super-resolution framework that combines image-language cognition with a latent diffusion prior. For a low-resolution input image (LR), a lightweight 4× SRResNet first produces an enhanced image used only for cognitive analysis. A CLIP image encoder and a learned cognitive adapter then produce a multi-token cognitive embedding containing semantic and appearance information. This embedding conditions a pretrained Stable Diffusion model to generate a reference image (GR) that is intended to be both semantically aligned with LR and visually similar in color and texture. Finally, a denoising U-Net reconstructs the super-resolved output (SR) using three conditions simultaneously: the LR image, the cognitive embedding, and GR. The cognitive embedding activates implicit diffusion priors, while GR provides an explicit image prior for high-frequency texture restoration.

  2. Knowl 2 — Multi-token cognitive encoder and supervision

    model/method

    The cognitive encoder converts CLIP image features from a preprocessed LR image into a language-compatible, multi-token embedding. Let BB be batch size, TiT_i the number of CLIP image tokens, CiC_i their channel dimension, TlT_l the number of CLIP language tokens, and ClC_l their channel dimension. The image and language embeddings are I∈RB×Ti×CiI\in\mathbb{R}^{B\times T_i\times C_i} and L∈RB×Tl×ClL\in\mathbb{R}^{B\times T_l\times C_l}, respectively. The language embedding is obtained from a caption generated from the corresponding HR image with BLIP2 during training. A cognitive adapter inspired by Q-Former uses Te≤TlT_e\leq T_l learnable queries to interact with spatial image tokens and produces E∈RB×Te×ClE\in\mathbb{R}^{B\times T_e\times C_l}. Rather than supervising only the CLIP class token, the adapter is supervised by a contiguous sequence of language tokens ending at the class token. If that sequence contains fewer than TeT_e tokens, the class token is repeated as padding. With tclst_{\mathrm{cls}} denoting the class-token position and inclusive token indexing, the target sequence is

    L′={Pad⁡Te ⁣(L0:tcls,Ltcls),tcls+1<Te,Ltcls−Te+1:tcls,tcls+1≥Te,L' = \begin{cases} \operatorname{Pad}_{T_e}\!\left(L_{0:t_{\mathrm{cls}}},L_{t_{\mathrm{cls}}}\right), & t_{\mathrm{cls}}+1<T_e,\\ L_{t_{\mathrm{cls}}-T_e+1:t_{\mathrm{cls}}}, & t_{\mathrm{cls}}+1\geq T_e, \end{cases}

    where Pad⁡Te\operatorname{Pad}_{T_e} extends the token sequence to length TeT_e by appending copies of the class token. The cognitive encoder is trained with the squared feature-matching loss

    LCE=∥E−L′∥22.\mathcal{L}_{\mathrm{CE}}=\lVert E-L'\rVert_2^2.

    Using multiple tokens reduces the semantic and texture bias observed when only the class token is aligned; the adapter also retains fine-grained visual information that direct captioning of degraded LR images may lose.

  3. Knowl 3 — Diffusion-generated reference prior

    model/method

    CoSeR uses the cognitive embedding EE as conditioning for a pretrained Stable Diffusion model to generate an explicit reference image without adding parameters to the diffusion generator. Because EE combines language-compatible semantics with visual appearance information, the generated reference is designed to preserve the LR image's subject, colors, and textures while supplying high-definition details. A pretrained VQGAN encodes both LR and reference images into latent codes. A single shared ControlNet-style U-Net encoder then extracts four multi-scale control features for the LR image, {Xi}i=14\{X_i\}_{i=1}^{4}, and four corresponding features for the reference image, {Ri}i=14\{R_i\}_{i=1}^{4}. Sharing one control encoder for both inputs reduces parameters while providing separate LR and reference controls. When the reference is generated automatically by Stable Diffusion, its latent code can be sent directly to the control encoder, avoiding an additional decode-and-re-encode operation.

  4. Knowl 4 — All-in-Attention conditional fusion

    model/method

    The All-in-Attention (AiA) module integrates LR fidelity, reference-image detail, and cognitive semantics inside the denoising U-Net. At each of four control scales, let ZZ denote the current U-Net feature, XiX_i the LR control, RiR_i the reference control, and QQ, KK, and VV denote attention queries, keys, and values. AiA augments the original Stable Diffusion attention block with trainable LR-attention and reference-attention branches while keeping the original self-attention and cross-attention branches frozen. The LR-attention branch uses Q=ZQ=Z and (K,V)(K,V) from XiX_i, directly enforcing consistency with LR content. The reference-attention branch uses Q=XiQ=X_i and (K,V)(K,V) from RiR_i, allowing LR-conditioned retrieval of corresponding reference details. The original cross-attention branch uses the cognitive embedding EE as its key and value source, while the U-Net feature supplies its query. These branches are applied throughout the middle and decoder portions of the denoising U-Net. CoSeR also adds one-hot attention to select the most relevant reference feature and reduce the blurring that can arise from conventional reference attention.

  5. Knowl 5 — Training and inference configuration

    experimental setup

    CoSeR is built on Stable Diffusion 2.1-base and trained at 512×512512\times512 resolution. The cognitive encoder is trained first with LCE\mathcal{L}_{\mathrm{CE}} and then frozen while the super-resolution model is trained. The SR model uses batch size 192 for 20,000 optimization steps on 8 V100 GPUs, with Adam and learning rate 5×10−55\times10^{-5}. ControlNet is initialized from Stable Diffusion weights, and the new LR-attention and reference-attention branches are initialized from self-attention weights. Inference uses DDPM sampling with 200 timesteps and classifier-free guidance with scale 3 for the cognitive condition. A pretrained feature-wrapping module is integrated with the VQGAN decoder to balance perceptual realism and fidelity.

  6. Knowl 6 — Training, test, and evaluation protocol

    experimental setup

    The model is trained on more than 900,000 ImageNet HR images at 512×512512\times512 resolution. Real-ESRGAN degradation produces paired LR images, and BLIP2 generates three descriptive captions per HR image; captions with CLIP scores below 0.28 are discarded. The authors construct a non-overlapping ImageNet Test2000 set containing 2,000 LR-HR pairs, selecting two images from each category for diversity and balance. RealSR and DRealSR provide additional real-world evaluation; their LR images are resized so the shorter side is 128 pixels and then center-cropped to 128×128128\times128, matching the training LR resolution. CoSeR is compared with RealSR, Real-ESRGAN+, BSRGAN, DASR, FeMaSR, LDM, and StableSR, with comparison models retrained on the same ImageNet training set where applicable. Evaluation uses FID, DISTS, and LPIPS for perceptual distance; CLIP-Score for semantic agreement with HR images; and MANIQA and MUSIQ for no-reference image quality.

  7. Knowl 7 — Benchmark performance across synthetic and real-world data

    data/table

    The quantitative comparison below evaluates perceptual distance, semantic accuracy, and no-reference quality. Lower is better for FID, DISTS, and LPIPS; higher is better for CLIP-Score, MANIQA, and MUSIQ. CoSeR obtains the best result on nearly every dataset-metric pair, although FeMaSR has a higher MUSIQ score on ImageNet Test2000. Relative to the second-best FID, the paper reports improvements of 13.8% on ImageNet Test2000, 3.8% on RealSR, and 4.7% on DRealSR.

    Dataset Metric RealSR Real-ESRGAN+ BSRGAN DASR FeMaSR LDM StableSR CoSeR
    ImageNet Test2000 FID↓\downarrow 86.36 32.68 41.11 39.15 31.25 34.54 22.53 19.41
    ImageNet Test2000 DISTS↓\downarrow 0.2649 0.1739 0.1946 0.1931 0.1597 0.1664 0.1527 0.1482
    ImageNet Test2000 LPIPS↓\downarrow 0.4519 0.2943 0.3381 0.3346 0.3027 0.3289 0.2871 0.2863
    ImageNet Test2000 CLIP-Score↑\uparrow 0.6242 0.8132 0.7719 0.7838 0.8253 0.8119 0.8622 0.8755
    ImageNet Test2000 MANIQA↑\uparrow 0.0796 0.1370 0.1115 0.0914 0.1936 0.1830 0.1556 0.2133
    ImageNet Test2000 MUSIQ↑\uparrow 50.18 57.52 52.33 48.98 67.20 64.15 60.20 65.51
    RealSR FID↓\downarrow 157.85 87.00 111.03 107.38 91.45 92.43 84.06 80.82
    RealSR DISTS↓\downarrow 0.2529 0.2028 0.2545 0.2171 0.2131 0.2055 0.1867 0.1826
    RealSR LPIPS↓\downarrow 0.3672 0.2803 0.3224 0.3056 0.2683 0.2924 0.2536 0.2438
    RealSR CLIP-Score↑\uparrow 0.7458 0.8345 0.8074 0.8332 0.8108 0.8330 0.8517 0.8545
    RealSR MANIQA↑\uparrow 0.1474 0.1776 0.1696 0.1803 0.2033 0.1986 0.2144 0.2522
    RealSR MUSIQ↑\uparrow 60.40 61.90 60.82 60.90 66.47 67.27 67.08 70.29
    DRealSR FID↓\downarrow 148.58 74.72 107.76 96.42 86.81 87.16 75.83 71.22
    DRealSR DISTS↓\downarrow 0.2673 0.2216 0.2238 0.2345 0.2231 0.2179 0.2048 0.1977
    DRealSR LPIPS↓\downarrow 0.4212 0.3239 0.3972 0.3534 0.2981 0.3258 0.2920 0.2702
    DRealSR CLIP-Score↑\uparrow 0.7360 0.8504 0.8157 0.8510 0.8332 0.8459 0.8681 0.8766
    DRealSR MANIQA↑\uparrow 0.1090 0.1742 0.1491 0.1739 0.1998 0.1890 0.2241 0.2575
    DRealSR MUSIQ↑\uparrow 54.28 62.80 57.72 62.14 66.57 67.03 68.27 70.18
  8. Knowl 8 — Cognitive and reference-prior ablations

    empirical result

    On ImageNet Test2000, the multi-token cognitive encoder outperforms both a class-token-only encoder and removing cognitive information. The cognitive comparison reports FID 20.27, CLIP-Score 0.8674, and generated-reference Gen-score 0.5953 for the multi-token encoder; FID 21.35, CLIP-Score 0.8628, and Gen-score 0.4881 for the class-token encoder; and FID 23.18 and CLIP-Score 0.8484 without a cognitive encoder. Gen-score is the CLIP-Score between a generated reference and its ground-truth HR image, so the higher multi-token score indicates less cognitive bias in reference generation.

    Explicit reference guidance also improves the final SR output. With a generated reference, the model obtains FID 19.80, MUSIQ 64.21, and MANIQA 0.2107; with a real ImageNet reference, it obtains FID 19.72, MUSIQ 63.49, and MANIQA 0.2056; without any reference, it obtains FID 20.27, MUSIQ 61.82, and MANIQA 0.1874. Thus, generated references provide quality gains comparable to real references while avoiding manual reference selection. These ablations remove the feature-wrapping module so that the component effects are evaluated directly.

  9. Knowl 9 — All-in-Attention improves fidelity

    empirical result

    Replacing the spatial feature transform used in StableSR with AiA improves fidelity on ImageNet Test2000 when feature wrapping is removed. The AiA model obtains FID 20.27, DISTS 0.1502, and LPIPS 0.3076, whereas the SFT variant obtains FID 21.50, DISTS 0.1530, and LPIPS 0.3101. The FID reduction is 5.7% relative to SFT, and the lower DISTS and LPIPS values indicate improved perceptual agreement with the input-grounded HR targets.

  10. Knowl 10 — Qualitative recovery and human preference

    empirical result

    Qualitative comparisons show that CoSeR produces clearer fur and facial features in animal images, more realistic anemone tentacles and succulent leaves, and semantic details that are nearly absent from LR inputs. In the reported examples, CoSeR is the only method that restores a dhole's eyes and the sand inside an hourglass. A user study evaluated 20 real-world LR images collected online or with mobile phones. Twenty-three subjects selected the visually best result among outputs from Real-ESRGAN+, FeMaSR, StableSR, and CoSeR, producing 20×23=46020\times23=460 votes. Approximately 80% of participants selected CoSeR as having the best visual effect.

Coverage note — No substantial contributed material was deliberately omitted; supplementary implementation details, pixel metrics, and reference-only background were excluded because they do not add a comparably load-bearing contribution.

References

  1. 1.Adrian Bulat, Jing Yang, and Georgios Tzimiropoulos. To learn image super-resolution, use a gan to learn how to do image degradation first. In ECCV, pages 185–200, 2018. 2
  2. 2.Jianrui Cai, Hui Zeng, Hongwei Yong, Zisheng Cao, and Lei Zhang. Toward real-world single image super-resolution: A new benchmark and a new model. In ICCV, pages 3086–3095, 2019. 2, 6
  3. 3.Jiezhang Cao, Jingyun Liang, Kai Zhang, Yawei Li, Yulun Zhang, Wenguan Wang, and Luc Van Gool. Reference-based image super-resolution with deformable attention transformer. In ECCV, pages 325–342. Springer, 2022. 2, 3
  4. 4.Kelvin CK Chan, Xintao Wang, Xiangyu Xu, Jinwei Gu, and Chen Change Loy. Glean: Generative latent bank for large-factor image super-resolution. In CVPR, pages 14245–14254, 2021. 2, 3
  5. 5.Chang Chen, Zhiwei Xiong, Xinmei Tian, Zheng-Jun Zha, and Feng Wu. Camera lens super-resolution. In CVPR, pages 1652–1660, 2019. 2
  6. 6.Chang Chen, Zhiwei Xiong, Xinmei Tian, Zheng-Jun Zha, and Feng Wu. Camera lens super-resolution. In CVPR, pages 1652–1660, 2019. 2
  7. 7.Chaofeng Chen, Xinyu Shi, Yipeng Qin, Xiaoming Li, Xiaoguang Han, Tao Yang, and Shihui Guo. Real-world blind super-resolution via feature matching with implicit high-resolution priors. In ACMMM, pages 1329–1338, 2022. 6
  8. 8.Jin Chen, Jun Chen, Zheng Wang, Chao Liang, and Chia-Wen Lin. Identity-aware face super-resolution for low-resolution face recognition. SPL, 27:645–649, 2020. 2
  9. 9.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In CVPR, pages 248–255. Ieee, 2009. 6
  10. 10.Prafulla Dhariwal and Alexander Nichol. Diffusion models beat gans on image synthesis. NeurIPS, 34:8780–8794, 2021. 3
  11. 11.Keyan Ding, Kede Ma, Shiqi Wang, and Eero P Simoncelli. Image quality assessment: Unifying structure and texture similarity. TPAMI, 44(5):2567–2581, 2020. 6
  12. 12.Patrick Esser, Robin Rombach, and Bjorn Ommer. Taming transformers for high-resolution image synthesis. In CVPR, pages 12873–12883, 2021. 5
  13. 13.Ben Fei, Zhaoyang Lyu, Liang Pan, Junzhe Zhang, Weidong Yang, Tianyue Luo, Bo Zhang, and Bo Dai. Generative diffusion prior for unified image restoration and enhancement. In CVPR, pages 9935–9946, 2023. 2, 3
  14. 14.Jinjin Gu, Yujun Shen, and Bolei Zhou. Image processing using multi-code gan prior. In CVPR, pages 3012–3021, 2020. 2, 3
  15. 15.Bahadir K Gunturk, Aziz Umit Batur, Yucel Altunbasak, Monson H Hayes, and Russell M Mersereau. Eigenface-domain super-resolution for face recognition. TIP, 12(5):597–606, 2003. 2
  16. 16.Muhammad Haris, Greg Shakhnarovich, and Norimichi Ukita. Task-driven super resolution: Object detection in low-resolution images. In ICONIP, pages 387–395. Springer, 2021. 2
  17. 17.Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. Gans trained by a two time-scale update rule converge to a local nash equilibrium. NeurIPS, 30, 2017. 6
  18. 18.Jonathan Ho and Tim Salimans. Classifier-free diffusion guidance. arXiv preprint arXiv:2207.12598, 2022. 6
  19. 19.Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. NeurIPS, 33:6840–6851, 2020. 3, 5
  20. 20.Lianghua Huang, Di Chen, Yu Liu, Yujun Shen, Deli Zhao, and Jingren Zhou. Composer: Creative and controllable image synthesis with composable conditions. arXiv preprint arXiv:2302.09778, 2023. 2
  21. 21.Xiaozhong Ji, Yun Cao, Ying Tai, Chengjie Wang, Jilin Li, and Feiyue Huang. Real-world super-resolution via kernel estimation and noise injection. In CVPRW, pages 466–467, 2020. 6
  22. 22.Yuming Jiang, Kelvin CK Chan, Xintao Wang, Chen Change Loy, and Ziwei Liu. Robust reference-based super-resolution via c2-matching. In CVPR, pages 2103–2112, 2021. 2
  23. 23.Bahjat Kawar, Michael Elad, Stefano Ermon, and Jiaming Song. Denoising diffusion restoration models. NeurIPS, 35:23593–23606, 2022. 2, 3
  24. 24.Junjie Ke, Qifei Wang, Yilin Wang, Peyman Milanfar, and Feng Yang. Musiq: Multi-scale image quality transformer. In ICCV, pages 5148–5157, 2021. 6
  25. 25.Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014. 5
  26. 26.Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C Berg, Wan-Yen Lo, et al. Segment anything. ICCV, 2023. 2
  27. 27.Christian Ledig, Lucas Theis, Ferenc Huszar, Jose Caballero, Andrew Cunningham, Alejandro Acosta, Andrew Aitken, Alykhan Tejani, Johannes Totz, Zehan Wang, et al. Photorealistic single image super-resolution using a generative adversarial network. In CVPR, pages 4681–4690, 2017. 2, 3
  28. 28.Guangyuan Li, Wei Xing, Lei Zhao, Zehua Lan, Jiakai Sun, Zhanjie Zhang, Quanwei Zhang, Huaizhong Lin, and Zhijie Lin. Self-reference image super-resolution via pre-trained diffusion large model and window adjustable transformer. In ACMMM, pages 7981–7992, 2023. 3
  29. 29.Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. arXiv preprint arXiv:2301.12597, 2023. 3, 4, 6
  30. 30.Wenbo Li, Kun Zhou, Lu Qi, Liying Lu, and Jiangbo Lu. Best-buddy gans for highly detailed image super-resolution. In AAAI, pages 1412–1420, 2022. 3
  31. 31.Yu-Jhe Li, Shawn Hunt, Jinhyung Park, Matthew O’Toole, and Kris Kitani. Azimuth super-resolution for fmcw radar in autonomous driving. In CVPR, pages 17504–17513, 2023. 2
  32. 32.Jingyun Liang, Jiezhang Cao, Guolei Sun, Kai Zhang, Luc Van Gool, and Radu Timofte. Swinir: Image restoration using swin transformer. In ICCV, pages 1833–1844, 2021. 2
  33. 33.Jie Liang, Hui Zeng, and Lei Zhang. Efficient and degradation-adaptive network for real-world image super-resolution. In ECCV, pages 574–591. Springer, 2022. 6
  34. 34.Jie Liang, Hui Zeng, and Lei Zhang. Details or artifacts: A locally discriminative learning approach to realistic image super-resolution. In CVPR, pages 5657–5666, 2022. 3
  35. 35.Xinqi Lin, Jingwen He, Ziyan Chen, Zhaoyang Lyu, Ben Fei, Bo Dai, Wanli Ouyang, Yu Qiao, and Chao Dong. Diffbir: Towards blind image restoration with generative diffusion prior. arXiv preprint arXiv:2308.15070, 2023. 2, 3
  36. 36.Anran Liu, Yihao Liu, Jinjin Gu, Yu Qiao, and Chao Dong. Blind image super-resolution: A survey and beyond. TPAMI, 45(5):5461–5480, 2022. 2
  37. 37.Liying Lu, Wenbo Li, Xin Tao, Jiangbo Lu, and Jiaya Jia. Masa-sr: Matching acceleration and spatial adaptation for reference-based image super-resolution. In CVPR, pages 6368–6377, 2021. 2, 3
  38. 38.Ziwei Luo, Fredrik K Gustafsson, Zheng Zhao, Jens Sjolund, and Thomas Sch{"o}n. Controlling vision-language models for universal image restoration. arXiv preprint arXiv:2310.01018, 2023. 3, 8
  39. 39.Yiyang Ma, Huan Yang, Wenjing Wang, Jianlong Fu, and Jiaying Liu. Unified multi-modal latent diffusion for joint subject and text conditional image generation. arXiv preprint arXiv:2303.09319, 2023. 3
  40. 40.Shunta Maeda. Unpaired image super-resolution using pseudo-supervision. In CVPR, pages 291–300, 2020. 2
  41. 41.Sachit Menon, Alexandru Damian, Shijia Hu, Nikhil Ravi, and Cynthia Rudin. Pulse: Self-supervised photo upsampling via latent space exploration of generative models. In CVPR, pages 2437–2445, 2020. 2, 3
  42. 42.Chong Mou, Xintao Wang, Liangbin Xie, Jian Zhang, Zhonggang Qi, Ying Shan, and Xiaohu Qie. T2i-adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models. arXiv preprint arXiv:2302.08453, 2023. 2
  43. 43.Kamyar Nazeri, Eric Ng, Tony Joseph, Faisal Z Qureshi, and Mehran Ebrahimi. Edgeconnect: Generative image inpainting with adversarial edge learning. arXiv preprint arXiv:1901.00212, 2019. 2
  44. 44.Xingang Pan, Xiaohang Zhan, Bo Dai, Dahua Lin, Chen Change Loy, and Ping Luo. Exploiting deep generative prior for versatile image restoration and manipulation. TPAMI, 44(11):7474–7489, 2021. 2, 3
  45. 45.Dustin Podell, Zion English, Kyle Lacey, Andreas Blattmann, Tim Dockhorn, Jonas Muller, Joe Penna, and Robin Rombach. Sdxl: Improving latent diffusion models for high-resolution image synthesis. arXiv preprint arXiv:2307.01952, 2023. 2, 3
  46. 46.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In ICML, pages 8748–8763. PMLR, 2021. 3, 4, 6
  47. 47.Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. Hierarchical text-conditional image generation with clip latents. arXiv preprint arXiv:2204.06125, 1(2):3, 2022. 2, 3
  48. 48.Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bjorn Ommer. High-resolution image synthesis with latent diffusion models. In CVPR, pages 10684–10695, 2022. 3, 6
  49. 49.Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily L Denton, Kamyar Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Salimans, et al. Photorealistic text-to-image diffusion models with deep language understanding. NeurIPS, 35:36479–36494, 2022. 2, 3
  50. 50.Chitwan Saharia, Jonathan Ho, William Chan, Tim Salimans, David J Fleet, and Mohammad Norouzi. Image super-resolution via iterative refinement. TPAMI, 45(4):4713–4726, 2022. 3
  51. 51.Shuwei Shi, Qingyan Bai, Mingdeng Cao, Weihao Xia, Jiahao Wang, Yifan Chen, and Yujiu Yang. Region-adaptive deformable network for image quality assessment. In CVPRW, pages 324–333, 2021. 6
  52. 52.Jiaming Song, Chenlin Meng, and Stefano Ermon. Denoising diffusion implicit models. arXiv preprint arXiv:2010.02502, 2020. 3
  53. 53.Jianyi Wang, Zongsheng Yue, Shangchen Zhou, Kelvin CK Chan, and Chen Change Loy. Exploiting diffusion prior for real-world image super-resolution. arXiv preprint arXiv:2305.07015, 2023. 2, 3, 5, 6, 8
  54. 54.Li Wang, Dong Li, Yousong Zhu, Lu Tian, and Yi Shan. Dual super-resolution learning for semantic segmentation. In CVPR, pages 3774–3783, 2020. 2
  55. 55.Ruoxi Wang, Dandan Zhang, Qingbiao Li, Xiao-Yun Zhou, and Benny Lo. Real-time surgical environment enhancement for robot-assisted minimally invasive surgery based on super-resolution. In ICRA, pages 3434–3440. IEEE, 2021. 2
  56. 56.Xintao Wang, Ke Yu, Chao Dong, and Chen Change Loy. Recovering realistic texture in image super-resolution by deep spatial feature transform. In CVPR, pages 606–615, 2018. 2, 8
  57. 57.Xintao Wang, Ke Yu, Shixiang Wu, Jinjin Gu, Yihao Liu, Chao Dong, Yu Qiao, and Chen Change Loy. Esrgan: Enhanced super-resolution generative adversarial networks. In ECCVW, pages 0–0, 2018. 2
  58. 58.Xintao Wang, Yu Li, Honglun Zhang, and Ying Shan. Towards real-world blind face restoration with generative facial prior. In CVPR, pages 9168–9178, 2021. 2, 3
  59. 59.Xintao Wang, Liangbin Xie, Chao Dong, and Ying Shan. Real-esrgan: Training real-world blind super-resolution with pure synthetic data. In ICCV, pages 1905–1914, 2021. 2, 3, 6
  60. 60.Yinhuai Wang, Jiwen Yu, and Jian Zhang. Zero-shot image restoration using denoising diffusion null-space model. arXiv preprint arXiv:2212.00490, 2022. 2, 3
  61. 61.Pengxu Wei, Ziwei Xie, Hannan Lu, Zongyuan Zhan, Qixiang Ye, Wangmeng Zuo, and Liang Lin. Component divide-and-conquer for real-world image super-resolution. In ECCV, pages 101–117. Springer, 2020. 2, 6
  62. 62.Yunxuan Wei, Shuhang Gu, Yawei Li, Radu Timofte, Longcun Jin, and Hengjie Song. Unsupervised real-world image super resolution via domain-distance aware training. In CVPR, pages 13385–13394, 2021. 2
  63. 63.Bin Xia, Yapeng Tian, Yucheng Hang, Wenming Yang, Qingmin Liao, and Jie Zhou. Coarse-to-fine embedded patchmatch and multi-scale dynamic aggregation for reference-based super-resolution. In AAAI, pages 2768–2776, 2022. 2, 3
  64. 64.Liangbin Xie, Xintao Wang, Xiangyu Chen, Gen Li, Ying Shan, Jiantao Zhou, and Chao Dong. Desra: Detect and delete the artifacts of gan-based real-world super-resolution models. In ICML, pages 38204–38226. PMLR, 2023. 3
  65. 65.Fuzhi Yang, Huan Yang, Jianlong Fu, Hongtao Lu, and Baining Guo. Learning texture transformer network for image super-resolution. In CVPR, pages 5791–5800, 2020. 2
  66. 66.Fuzhi Yang, Huan Yang, Jianlong Fu, Hongtao Lu, and Baining Guo. Learning texture transformer network for image super-resolution. In CVPR, pages 5791–5800, 2020. 2, 3, 5
  67. 67.Sidi Yang, Tianhe Wu, Shuwei Shi, Shanshan Lao, Yuan Gong, Mingdeng Cao, Jiahao Wang, and Yujiu Yang. Maniqa: Multi-dimension attention network for no-reference image quality assessment. In CVPR, pages 1191–1200, 2022. 6
  68. 68.Tao Yang, Peiran Ren, Xuansong Xie, and Lei Zhang. Gan prior embedded network for blind face restoration in the wild. In CVPR, pages 672–681, 2021. 2, 3
  69. 69.Tao Yang, Peiran Ren, Xuansong Xie, and Lei Zhang. Pixel-aware stable diffusion for realistic image super-resolution and personalized stylization. arXiv preprint arXiv:2308.14469, 2023. 2, 3
  70. 70.Tao Yang, Peiran Ren, Lei Zhang, et al. Synthesizing realistic image restoration training pairs: A diffusion approach. arXiv preprint arXiv:2303.06994, 2023. 2
  71. 71.Yuan Yuan, Siyuan Liu, Jiawei Zhang, Yongbing Zhang, Chao Dong, and Liang Lin. Unsupervised image super-resolution using cycle-in-cycle generative adversarial networks. In CVPRW, pages 701–710, 2018. 2
  72. 72.Zongsheng Yue and Chen Change Loy. Difface: Blind face restoration with diffused error contraction. arXiv preprint arXiv:2212.06512, 2022. 2, 3
  73. 73.Kai Zhang, Jingyun Liang, Luc Van Gool, and Radu Timofte. Designing a practical degradation model for deep blind image super-resolution. In ICCV, pages 4791–4800, 2021. 2, 3, 6
  74. 74.Lvmin Zhang, Anyi Rao, and Maneesh Agrawala. Adding conditional control to text-to-image diffusion models. In ICCV, pages 3836–3847, 2023. 2, 5, 6
  75. 75.Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shechtman, and Oliver Wang. The unreasonable effectiveness of deep features as a perceptual metric. In CVPR, pages 586–595, 2018. 6
  76. 76.Yulun Zhang, Kunpeng Li, Kai Li, Lichen Wang, Bineng Zhong, and Yun Fu. Image super-resolution using very deep residual channel attention networks. In ECCV, pages 286–301, 2018. 2
  77. 77.Zhifei Zhang, Zhaowen Wang, Zhe Lin, and Hairong Qi. Image super-resolution by neural texture transfer. In CVPR, pages 7982–7991, 2019. 2, 3
  78. 78.Shihao Zhao, Dongdong Chen, Yen-Chun Chen, Jianmin Bao, Shaozhe Hao, Lu Yuan, and Kwan-Yee K Wong. Uni-controlnet: All-in-one control to text-to-image diffusion models. arXiv preprint arXiv:2305.16322, 2023. 2
  79. 79.Haitian Zheng, Mengqi Ji, Haoqian Wang, Yebin Liu, and Lu Fang. Crossnet: An end-to-end reference-based super resolution network using cross-scale warping. In ECCV, pages 88–104, 2018. 2, 3
  80. 80.Shangchen Zhou, Jiawei Zhang, Wangmeng Zuo, and Chen Change Loy. Cross-scale internal graph neural network for image super-resolution. NeurIPS, 33:3499–3509, 2020. 2

Citation

MLA
Sun, H., et al. “CoSeR: Bridging Image and Language for Cognitive Super-Resolution”. arXiv, 2023, http://arxiv.org/abs/2311.16512v4.
APA
Sun, H., Li, W., Liu, J., Chen, H., Pei, R., Zou, X., Yan, Y., & Yang, Y. (2023). CoSeR: Bridging Image and Language for Cognitive Super-Resolution. arXiv. http://arxiv.org/abs/2311.16512v4
Chicago
Sun, H., W. Li, J. Liu, et al. 2023. “CoSeR: Bridging Image and Language for Cognitive Super-Resolution”. arXiv. http://arxiv.org/abs/2311.16512v4.
Harvard
Sun, H. et al. (2023) “CoSeR: Bridging Image and Language for Cognitive Super-Resolution”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2311.16512v4.
Vancouver
1. Sun H, Li W, Liu J, Chen H, Pei R, Zou X, Yan Y, Yang Y (2023) CoSeR: Bridging Image and Language for Cognitive Super-Resolution. arXiv

BibTeX

@article{sun2023coser,
  title = {CoSeR: Bridging Image and Language for Cognitive Super-Resolution},
  author = {Sun, Haoze and Li, Wenbo and Liu, Jianzhuang and Chen, Haoyu and Pei, Renjing and Zou, Xueyi and Yan, Youliang and Yang, Yujiu},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2311.16512v4},
  eprint = {2311.16512}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE