MUSIQ: Multi-scale Image Quality Transformer

Junjie KeQifei WangYilin WangPeyman MilanfarFeng Yang

article2021ICCV1,656 citations

Proposes a multi-scale vision Transformer that evaluates full-resolution images of varying sizes and aspect ratios without quality-degrading cropping or resizing, achieving state-of-the-art accuracy on standard image quality assessment benchmarks.

Listen

Assessing the perceptual quality of digital images is vital for optimizing consumer visual experiences across digital platforms, photography, and video delivery. Standard deep learning approaches rely on convolutional neural networks that require images to be cropped or resized to uniform square shapes during batch training. This artificial distortion alters image composition and degrades fine visual details, compromising the accuracy of quality predictions in real-world scenarios where images vary widely in resolution and aspect ratio.

To address this limitation, the article introduces the Multi-scale Image Quality Transformer (MUSIQ). The primary objective is to demonstrate an assessment model that can process full-size images at native resolutions and across diverse aspect ratios without destructive preprocessing, capturing image quality across multiple levels of granular detail.

The approach employs a patch-based Transformer architecture inspired by human vision. Instead of forcing images into fixed dimensions, MUSIQ extracts patches from the full-resolution image alongside aspect-ratio-preserving resized variants. To maintain spatial and multi-scale context within the sequence of patches, the model incorporates a lightweight five-layer convolutional patch encoder, a hash-based two-dimensional spatial embedding, and a scale embedding. The model was pre-trained on standard ImageNet data and evaluated across four large-scale benchmark datasets encompassing both technical quality and aesthetic assessments.

The evaluation yielded several key findings. First, MUSIQ established state-of-the-art performance across three major technical quality datasets—PaQ-2-PiQ, KonIQ-10k, and SPAQ—and matched leading methods on the aesthetic assessment benchmark AVA. Second, on the PaQ-2-PiQ test set containing high-resolution images exceeding 640 pixels, the model achieved a linear correlation of 0.739, outperforming previous approaches by a noticeable margin. Third, ablation analyses confirmed that preserving the native aspect ratio is critical, as models forced into square resizing failed to detect realistic quality degradation caused by distortion. Finally, multi-scale representation outperformed both single-scale inputs and simple model ensembles, as the architecture effectively focused on fine details in high-resolution patches while evaluating global composition in lower-resolution views.

These findings demonstrate that computer vision systems for image assessment no longer need to sacrifice image fidelity for computational convenience. Operating at approximately 27 million parameters, MUSIQ delivers state-of-the-art accuracy with computational complexity comparable to standard ResNet-50 baselines. This eliminates the latency and storage costs associated with multi-crop sampling or offline feature caching, providing an efficient, end-to-end framework suitable for production pipelines in content moderation, camera tuning, and media delivery.

Organizations handling variable-format visual content should consider adopting patch-based Transformer architectures that preserve native image properties rather than standard fixed-crop convolutional networks. Future efforts should explore integrating efficient Transformer variants, such as linear-complexity attention mechanisms, to further accelerate training and inference speed on ultra-high-resolution media.

Confidence in these findings is high, supported by rigorous cross-dataset benchmarking and consistent improvements over 10-run averages. Readers should note that for extremely large images, memory limits may require capping the maximum patch count during training, meaning very high-resolution images could experience slight truncation if patch limits are set too low.

Cover for MUSIQ: Multi-scale Image Quality Transformer

Abstract

Image quality assessment (IQA) is an important research topic for understanding and improving visual experience. The current state-of-the-art IQA methods are based on convolutional neural networks (CNNs). The performance of CNN-based models is often compromised by the fixed shape constraint in batch training. To accommodate this, the input images are usually resized and cropped to a fixed shape, causing image quality degradation. To address this, we design a multi-scale image quality Transformer (MUSIQ) to process native resolution images with varying sizes and aspect ratios. With a multi-scale image representation, our proposed method can capture image quality at different granularities. Furthermore, a novel hash-based 2D spatial embedding and a scale embedding is proposed to support the positional embedding in the multi-scale representation. Experimental results verify that our method can achieve state-of-the-art performance on multiple large scale IQA datasets such as PaQ-2-PiQ, SPAQ and KonIQ-10k.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Multi-scale Image Quality Transformer
  • 3.1 Overall Architecture
  • 3.2 Multi-scale Patch Embedding
  • 3.3 Hash-based 2D Spatial Embedding
  • 3.4 Scale Embedding
  • 3.5 Pre-training and Fine-tuning
  • 4 Experimental Results
  • 4.1 Datasets
  • 4.2 Implementation Details
  • 4.3 Comparing with the State-of-the-art (SOTA)
  • 4.4 Ablation Studies
  • 5 Conclusion
  • References
  • A Transformer Encoder
  • A.1 Transformer Encoder Structure
  • A.2 Multi-head Self-Attention (MSA)
  • A.3 Masked Self-Attention
  • A.4 Different Transformer Encoder Settings
  • B Additional Studies for HSE
  • B.1 Grid Size GG in HSE
  • B.2 Sinusoidal HSE v.s. Learnable HSE
  • B.3 Visualization of HSE with Different GG
  • C Effect of Patch Size
  • D The Maximum Number of Patches (ll) from Full-size Image
  • E KonIQ-10k More Results
  • F SPAQ Full-size Results
  • G Computation Complexity
  • H Multi-scale Attention Visualization

Knowls

  1. Knowl 1 — Multi-Scale Image Quality Transformer (MUSIQ) Architecture

    model/method

    The Multi-scale Image Quality Transformer (MUSIQ) is a patch-based Transformer framework designed for blind image quality assessment (IQA) that natively processes full-size images of arbitrary resolutions and aspect ratios without requiring aspect-ratio-distorting preprocessing.

    Given an input image, MUSIQ constructs a multi-scale representation consisting of the full-size native resolution image and KK aspect-ratio-preserving (ARP) resized variants. Each image in this multi-scale pyramid is partitioned into non-overlapping square patches of size P×PP \times P. Each patch is mapped into a DD-dimensional token representation using a lightweight shared 5-layer residual convolutional network (ResNet). To retain geometric structure and scale context across varying resolutions, two additive embedding components are added element-wise to each patch token:

    1. A Hash-based 2D Spatial Embedding (HSE), which maps the 2D coordinates of patches to a fixed G×GG \times G learnable grid, aligning spatially proximate regions across different scales.
    2. A Scale Embedding (SCE), which explicitly indicates the scale level k∈{0,…,K}k \in \{0, \dots, K\} from which the patch originated.

    A learnable classification token embedding ([CLS][\text{CLS}]) is prepended to the combined token sequence. The resulting sequence is processed by a Transformer encoder consisting of LL layers of Multi-Head Self-Attention (MSA) and Multi-Layer Perceptrons (MLP) with Layer Normalization and residual skip connections. The output state corresponding to the [CLS][\text{CLS}] token serves as the global quality representation and is passed to a fully connected prediction head to output the predicted quality score or score distribution.

  2. Knowl 2 — Hash-Based 2D Spatial Embedding (HSE)

    model/method

    Standard fixed-length 1D positional embeddings fail when input images have arbitrary resolutions and aspect ratios because the sequence length changes and 1D indexes do not correspond to fixed 2D locations or align across scales. The Hash-based 2D Spatial Embedding (HSE) addresses this by mapping patch positions onto a shared G×GG \times G learnable embedding grid T∈RG×G×DT \in \mathbb{R}^{G \times G \times D}, where GG is the grid resolution and DD is the token embedding dimension.

    For an image of resolution H×WH \times W partitioned into patches of size P×PP \times P, the patch located at row index i∈{0,…,HP−1}i \in \{0, \dots, \frac{H}{P}-1\} and column index j∈{0,…,WP−1}j \in \{0, \dots, \frac{W}{P}-1\} has spatial embedding coordinates (ti,tj)(t_i, t_j) computed by continuous hashing and rounding to the nearest integer grid cell:

    ti=round(i×GH/P),tj=round(j×GW/P)t_i = \text{round}\left(\frac{i \times G}{H / P}\right), \quad t_j = \text{round}\left(\frac{j \times G}{W / P}\right)

    The DD-dimensional spatial embedding vector Tti,tjT_{t_i, t_j} is retrieved and added element-wise to the patch token embedding.

    Because patch coordinates (i,j)(i, j) and image dimensions (H,W)(H, W) scale proportionally by the resizing factor αk\alpha_k across all aspect-ratio-preserving resized variants, patches covering the same physical region of the scene at different scales map to identical or adjacent cells in TT. This produces spatial alignment across scales while remaining computationally lightweight and non-intrusive to standard self-attention.

  3. Knowl 3 — Multi-Scale Image Representation and Scale Embedding (SCE)

    model/method

    To capture both fine-grained local textures and coarse global scene composition, MUSIQ represents an input image as a multi-scale pyramid consisting of the native resolution image I0∈RH×W×CI_0 \in \mathbb{R}^{H \times W \times C} and KK aspect-ratio-preserving (ARP) resized variants Ik∈Rhk×wk×CI_k \in \mathbb{R}^{h_k \times w_k \times C} for k∈{1,…,K}k \in \{1, \dots, K\}. Resized variants are downscaled using a Gaussian kernel with fixed maximum longer side lengths LkL_k:

    αk=Lkmax⁡(H,W),hk=round(αkH),wk=round(αkW)\alpha_k = \frac{L_k}{\max(H, W)}, \quad h_k = \text{round}(\alpha_k H), \quad w_k = \text{round}(\alpha_k W)

    where αk\alpha_k is the scale factor for scale kk. In the default MUSIQ configuration, K=2K=2 with L1=224L_1 = 224 and L2=384L_2 = 384.

    Each scale image is partitioned into P×PP \times P patches (P=32P=32). The number of patches extracted from native and resized images are N=HW/P2N = HW/P^2 and nk=hkwk/P2n_k = h_k w_k / P^2, respectively. Because the spatial hashing matrix TT is shared across scales, a learnable Scale Embedding matrix Q∈R(K+1)×DQ \in \mathbb{R}^{(K+1) \times D} is introduced to differentiate patch sources. The embedding vector Q0∈RDQ_0 \in \mathbb{R}^D is added element-wise to all patch tokens from the native image, while Qk∈RDQ_k \in \mathbb{R}^D is added element-wise to all patch tokens from resized variant kk.

  4. Knowl 4 — Masked Multi-Head Self-Attention for Variable Patch Sequences

    equation

    Because input images have varying resolutions and aspect ratios, the total number of patch tokens Ntotal=N+∑k=1KnkN_{\text{total}} = N + \sum_{k=1}^K n_k varies per sample. To support batch training on accelerators, the patch sequence for the native resolution image is padded or truncated to a fixed maximum length ll (e.g., l=512l=512), while each resized variant kk is padded to its upper bound mk=⌊Lk2/P2⌋m_k = \lfloor L_k^2 / P^2 \rfloor.

    To prevent zero-padded dummy tokens from influencing self-attention calculations, an attention mask matrix M∈RNseq×NseqM \in \mathbb{R}^{N_{\text{seq}} \times N_{\text{seq}}} is applied within each attention head:

    Mi,j={0,if token i and token j are both valid tokens−∞,if token i or token j is a padding tokenM_{i,j} = \begin{cases} 0, & \text{if token } i \text{ and token } j \text{ are both valid tokens} \\ -\infty, & \text{if token } i \text{ or token } j \text{ is a padding token} \end{cases}

    Given queries Q=zUqQ = z U_q, keys K=zUkK = z U_k, and values V=zUvV = z U_v generated from token sequence z∈RNseq×Dz \in \mathbb{R}^{N_{\text{seq}} \times D} via learnable projection matrices Uq,Uk,Uv∈RD×DhU_q, U_k, U_v \in \mathbb{R}^{D \times D_h}, the masked self-attention weight matrix AmA_m and single-head output SA(z)\text{SA}(z) are defined as:

    Am=softmax(QKT+MDh)A_m = \text{softmax}\left(\frac{Q K^T + M}{\sqrt{D_h}}\right)

    SA(z)=AmV\text{SA}(z) = A_m V

    Multi-head self-attention (extMSA ext{MSA}) concatenates the outputs of ss parallel heads and projects them using output matrix Um∈RsDh×DU_m \in \mathbb{R}^{s D_h \times D} with Dh=D/sD_h = D/s:

    MSA(z)=[SA1(z);… ;SAs(z)]Um\text{MSA}(z) = [\text{SA}_1(z); \dots; \text{SA}_s(z)] U_m

  5. Knowl 5 — Pre-training and Fine-tuning Optimization Protocol for MUSIQ

    experimental setup

    MUSIQ uses a two-stage training strategy:

    1. Pre-training: Pre-trained on ILSVRC-2012 ImageNet for 300 epochs using the Adam optimizer with β1=0.9\beta_1 = 0.9, β2=0.999\beta_2 = 0.999, weight decay 0.10.1, batch size 40964096, and cosine learning rate decay starting from 0.0010.001. Crucially, images are cropped to random sizes rather than resized to a fixed square shape, preserving native aspect ratios and priming the model for variable-resolution inputs. Data augmentations include RandAugment and Mixup.

    2. Fine-tuning: Fine-tuned on target IQA datasets using SGD with momentum and cosine learning rate decay. No cropping or resizing is performed during fine-tuning (only random horizontal flips are applied). Initial learning rates and epoch counts are:

      • PaQ-2-PiQ: LR 0.00020.0002, 1010 epochs, batch size 128128
      • KonIQ-10k: LR 0.00010.0001, 3030 epochs, batch size 9696
      • SPAQ: LR 0.00010.0001, 3030 epochs, batch size 128128
      • AVA: LR 0.120.12, 2020 epochs, batch size 512512

    For datasets labeled with a single Mean Opinion Score (MOS), an L1L_1 loss is minimized. For datasets with score distribution labels (AVA), the Earth Mover's Distance (EMD) loss with r=2r=2 is minimized:

    EMD(p,p^)=(1Nbins∑m=1Nbins∣CDFp(m)−CDFp^(m)∣r)1/r\text{EMD}(p, \hat{p}) = \left( \frac{1}{N_{\text{bins}}} \sum_{m=1}^{N_{\text{bins}}} \left| \text{CDF}_p(m) - \text{CDF}_{\hat{p}}(m) \right|^r \right)^{1/r}

    where pp is the ground-truth normalized distribution, p^\hat{p} is the predicted distribution, and CDFp(m)=∑i=1mpi\text{CDF}_p(m) = \sum_{i=1}^m p_i.

  6. Knowl 6 — Benchmark Evaluation on IQA Datasets

    data/table

    MUSIQ was evaluated across four large-scale IQA benchmarks: PaQ-2-PiQ, KonIQ-10k, and SPAQ for technical quality assessment, and AVA for aesthetic visual analysis. Performance was measured via Spearman Rank-Order Correlation Coefficient (SRCC) and Pearson Linear Correlation Coefficient (PLCC), averaged over 10 runs.

    Dataset PaQ-2-PiQ (Test) KonIQ-10k SPAQ AVA
    Metric SRCC PLCC SRCC PLCC SRCC PLCC SRCC PLCC
    BRISQUE 0.288 0.373 0.665 0.681 0.809 0.817 - -
    NIMA 0.583 0.639 - - - - 0.612 0.636
    DBCNN - - 0.875 0.884 0.911 0.915 - -
    BIQA (25 crops) - - 0.906 0.917 - - - -
    Fang et al. - - - - 0.908 0.909 - -
    Hosu et al. (20 crops) - - - - - - 0.756 0.757
    Ying et al. 0.601 0.685 - - - - - -
    MUSIQ-single 0.640 0.721 0.905 0.919 0.917 0.920 0.719 0.731
    MUSIQ (Multi-scale) 0.646 0.739 0.916 0.928 0.917 0.921 0.726 0.738

    MUSIQ consistently outperforms previous CNN-based methods across all technical IQA datasets and matches or exceeds multi-crop CNN ensembles on aesthetic quality (AVA MSE: 0.242 vs. 0.271 for AFDC+SPP), achieving superior quality representation in a single forward pass without crop sampling.

  7. Knowl 7 — Ablation on Aspect-Ratio Preservation vs. Square Resizing

    data/table

    To evaluate the degradation caused by standard square resizing in CNN- and Transformer-based IQA models, MUSIQ and baseline architectures were trained and evaluated on the AVA dataset using square inputs versus aspect-ratio-preserving (ARP) resizing.

    Method Input Resolution SRCC PLCC
    NIMA (Inception-v2) 224×224224 \times 224 square 0.612 0.636
    NIMA (ResNet50) 384×384384 \times 384 square 0.624 0.632
    ViT-Base 32 384×384384 \times 384 square 0.654 0.664
    ViT-Small 32 384×384384 \times 384 square 0.656 0.665
    MUSIQ square resizing (512,384,224)(512, 384, 224) 0.706 0.720
    MUSIQ ARP resizing (512,384,224)(512, 384, 224) 0.712 0.726
    MUSIQ ARP resizing (full,384,224)(\text{full}, 384, 224) 0.726 0.738

    Preserving aspect ratio improves performance over square resizing by 0.0060.006 SRCC on fixed scales, and retaining the full native resolution with ARP provides an additional gain of 0.0140.014 SRCC and 0.0120.012 PLCC. Furthermore, tests on images distorted to artificial aspect ratios demonstrate that MUSIQ correctly penalizes unnatural deformations, whereas models trained on square resized images are insensitive to aspect ratio distortions.

  8. Knowl 8 — Ablation of Multi-Scale Representation vs. Single Scale and Ensembles

    data/table

    The effect of multi-scale representation composition was evaluated on the PaQ-2-PiQ full-size test set by testing single resolutions, multi-scale combinations, and score averaging ensembles.

    Multi-scale Composition SRCC PLCC
    (224)(224) 0.600 0.667
    (384)(384) 0.618 0.695
    (512)(512) 0.620 0.691
    (384,224)(384, 224) 0.620 0.707
    (512,384,224)(512, 384, 224) 0.629 0.718
    (full)(\text{full}) 0.640 0.721
    (full,224)(\text{full}, 224) 0.643 0.726
    (full,384)(\text{full}, 384) 0.642 0.730
    (full,384,224)(\text{full}, 384, 224) 0.646 0.739
    Average ensemble of (full)(\text{full}), (224)(224), (384)(384) 0.640 0.710

    Adding ARP resized variants to the native full-size image consistently improves prediction accuracy over single-scale models (0.6460.646 vs. 0.6400.640 SRCC). Crucially, the multi-scale Transformer surpasses a naive average ensemble of separate models evaluated at individual scales (0.6400.640 SRCC, 0.7100.710 PLCC), demonstrating that cross-scale self-attention dynamically synthesizes complementary fine-grained and coarse-grained visual information.

  9. Knowl 9 — Ablation of Spatial and Scale Embeddings

    data/table

    Ablation experiments conducted on the AVA dataset isolate the contributions of the Hash-based 2D Spatial Embedding (HSE) and the Scale Embedding (SCE), as well as the HSE grid size GG and embedding parameterization.

    Spatial Embedding Configuration SRCC PLCC
    Without Spatial Embedding 0.704 0.716
    Fixed-length 1D Positional Embedding 0.707 0.722
    HSE (G=10G=10) 0.726 0.738
    Scale Embedding Configuration SRCC PLCC
    Without Scale Embedding (w/o SCE) 0.717 0.729
    With Scale Embedding (w/ SCE) 0.726 0.738

    Fixed-length 1D positional embeddings fail to capture aspect-ratio shifts and cannot align cross-scale coordinates. HSE yields a +0.019 gain in SRCC over 1D embeddings. Adding SCE provides an additional +0.009 SRCC improvement.

    Ablating grid size GG for learnable HSE reveals optimal performance around G=10G=10 (0.7260.726 SRCC / 0.7380.738 PLCC for G=10G=10 vs. 0.7200.720 / 0.7330.733 for G=5G=5 and 0.7220.722 / 0.7340.734 for G=20G=20), following the heuristic G×G×P×P≈H×WG \times G \times P \times P \approx H \times W. Learnable HSE grids consistently outperform sinusoidal fixed HSE grids (0.7260.726 vs. 0.7190.719 SRCC at G=10G=10).

  10. Knowl 10 — Ablation of Patch Encoder Architecture and Patch Size

    data/table

    Ablations on the patch token encoder module (evaluated on the PaQ-2-PiQ full-size test set) and patch size PP (evaluated on the AVA dataset) were conducted to establish optimal token generation settings.

    Patch Encoder Architecture Model Parameters SRCC PLCC
    Linear Projection 22M 0.634 0.714
    Simple Conv (7×77\times 7 followed by 3×33\times 3) 23M 0.639 0.726
    5-layer ResNet (Simple Conv + Residual Block) 27M 0.646 0.739
    Patch Size PP 16 32 48 64
    SRCC 0.715 0.726 0.713 0.705
    PLCC 0.729 0.738 0.727 0.719

    Using a 5-layer ResNet patch encoder outperforms linear projection by +0.012+0.012 SRCC and +0.025+0.025 PLCC with modest parameter overhead. Patch size P=32P=32 provides the optimal balance between spatial granularity and computational tractability across datasets.

Coverage note — None was omitted; all key contributions including the multi-scale architecture, hash-based spatial and scale embeddings, training procedures, benchmark evaluations, and ablation studies are covered.

References

  1. 1.Edward H Adelson, Charles H Anderson, James R Bergen, Peter J Burt, and Joan M Ogden. Pyramid methods in image processing. RCA engineer, 29(6):33–41, 1984. 2
  2. 2.Irwan Bello, Barret Zoph, Ashish Vaswani, Jonathon Shlens, and Quoc V Le. Attention augmented convolutional net-works. In Proceedings of the IEEE/CVF International Con-ference on Computer Vision, pages 3286–3295, 2019. 2, 3
  3. 3.Sebastian Bosse, Dominique Maniry, Klaus-Robert Muller, ¨ Thomas Wiegand, and Wojciech Samek. Deep neural net-works for no-reference and full-reference image quality as-sessment. IEEE Transactions on image processing, 27(1): 206–219, 2017. 6
  4. 4.Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. End-to-end object detection with transformers. In European Confer-ence on Computer Vision, pages 213–229. Springer, 2020. 1, 2
  5. 5.Hanting Chen, Yunhe Wang, Tianyu Guo, Chang Xu, Yiping Deng, Zhenhua Liu, Siwei Ma, Chunjing Xu, Chao Xu, and Wen Gao. Pre-trained image processing transformer. arXiv preprint arXiv:2012.00364, 2020.
  6. 6.Mark Chen, Alec Radford, Rewon Child, Jeffrey Wu, Hee-woo Jun, David Luan, and Ilya Sutskever. Generative pre-training from pixels. In Hal Daume III and Aarti Singh, ´ editors, Proceedings of the 37th International Conference on Machine Learning, volume 119 of Proceedings of Ma-chine Learning Research, pages 1691–1703. PMLR, 13–18 Jul 2020. 1, 2
  7. 7.Qiuyu Chen, Wei Zhang, Ning Zhou, Peng Lei, Yi Xu, Yu Zheng, and Jianping Fan. Adaptive fractional dilated con-volution network for image aesthetics assessment. In Pro-ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 14114–14123, 2020. 1, 2, 6, 7
  8. 8.Ekin D Cubuk, Barret Zoph, Jonathon Shlens, and Quoc V Le. Randaugment: Practical automated data augmenta-tion with a reduced search space. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, pages 702–703, 2020. 5
  9. 9.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition, pages 248–255. Ieee, 2009. 12
  10. 10.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT: pre-training of deep bidirectional trans-formers for language understanding. In Jill Burstein, Christy Doran, and Thamar Solorio, editors, Proceedings of the 2019 Conference of the North American Chapter of the Associa-tion for Computational Linguistics: Human Language Tech-nologies, NAACL-HLT 2019, Minneapolis, MN, USA, June 2-7, 2019, Volume 1 (Long and Short Papers), pages 4171– 4186. Association for Computational Linguistics, 2019. doi: 10.18653/v1/n19-1423. 2, 3, 11
  11. 11.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Syl-vain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition at scale. In International Conference on Learning Repre-sentations, 2021. URL https://openreview.net/forum?id=YicbFdNTTy. 1, 2, 3, 4, 7, 8, 11, 12, 14
  12. 12.Yuming Fang, Hanwei Zhu, Yan Zeng, Kede Ma, and Zhou Wang. Perceptual quality assessment of smartphone photog-raphy. In Proceedings of the IEEE/CVF Conference on Com-puter Vision and Pattern Recognition, pages 3677–3686, 2020. 1, 2, 5, 6, 13
  13. 13.Jonas Gehring, Michael Auli, David Grangier, Denis Yarats, and Yann N Dauphin. Convolutional sequence to sequence learning. In International Conference on Machine Learning, pages 1243–1252. PMLR, 2017. 3
  14. 14.Deepti Ghadiyaram and Alan C Bovik. Perceptual quality prediction on authentically distorted images using a bag of features approach. Journal of Vision, 17(1):32–32, 2017. 2, 6
  15. 15.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceed-ings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 770–778, 2016. 4
  16. 16.Vlad Hosu, Bastian Goldlucke, and Dietmar Saupe. Effective aesthetics prediction with multi-level spatially pooled fea-tures. In Proceedings of the IEEE/CVF Conference on Com-puter Vision and Pattern Recognition, pages 9375–9383, 2019. 1, 2, 6
  17. 17.Vlad Hosu, Hanhe Lin, Tamas Sziranyi, and Dietmar Saupe. Koniq-10k: An ecologically valid database for deep learning of blind image quality assessment. IEEE Transactions on Image Processing, 29:4041–4056, 2020. 1, 2, 5, 13
  18. 18.Le Kang, Peng Ye, Yi Li, and David Doermann. Convolu-tional neural networks for no-reference image quality assess-ment. In Proceedings of the IEEE conference on Computer Vision and Pattern Recognition, pages 1733–1740, 2014. 6
  19. 19.Jongyoo Kim and Sanghoon Lee. Fully deep blind image quality predictor. IEEE Journal of Selected Topics in Signal Processing, 11(1):206–220, 2016. 6
  20. 20.Shu Kong, Xiaohui Shen, Zhe Lin, Radomir Mech, and Charless Fowlkes. Photo aesthetics ranking network with attributes and content adaptation. In European Conference on Computer Vision, pages 662–679. Springer, 2016. 6
  21. 21.Dingquan Li, Tingting Jiang, Weisi Lin, and Ming Jiang. Which has better visual quality: The clear blue sky or a blurry animal? IEEE Transactions on Multimedia, 21(5): 1221–1234, 2018. 6
  22. 22.Tsung-Yi Lin, Piotr Dollar, Ross Girshick, Kaiming He, ´ Bharath Hariharan, and Serge Belongie. Feature pyramid networks for object detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 2117–2125, 2017. 2
  23. 23.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettle-moyer, and Veselin Stoyanov. Roberta: A robustly optimized BERT pretraining approach. CoRR, abs/1907.11692, 2019. URL http://arxiv.org/abs/1907.11692. 2
  24. 24.Shuang Ma, Jing Liu, and Chang Wen Chen. A-lamp: Adap-tive layout-aware multi-patch deep convolutional neural network for photo aesthetic assessment. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recogni-tion, pages 4535–4544, 2017. 2, 6
  25. 25.Long Mai, Hailin Jin, and Feng Liu. Composition-preserving deep photo aesthetics assessment. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recog-nition, pages 497–506, 2016. 1, 2, 6
  26. 26.Anish Mittal, Anush Krishna Moorthy, and Alan Conrad Bovik. No-reference image quality assessment in the spatial domain. IEEE Transactions on image processing, 21(12): 4695–4708, 2012. 2, 6
  27. 27.Anish Mittal, Rajiv Soundararajan, and Alan C Bovik. Mak-ing a “completely blind” image quality analyzer. IEEE Sig-nal processing letters, 20(3):209–212, 2012. 6
  28. 28.Anush Krishna Moorthy and Alan Conrad Bovik. Blind im-age quality assessment: From natural scene statistics to per-ceptual quality. IEEE transactions on Image Processing, 20 (12):3350–3364, 2011. 2, 6
  29. 29.Naila Murray and Albert Gordo. A deep architecture for uni-fied aesthetic prediction. arXiv preprint arXiv:1708.04890, 2017. 6
  30. 30.Naila Murray, Luca Marchesotti, and Florent Perronnin. Ava: A large-scale database for aesthetic visual analysis. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 2408–2415. IEEE, 2012. 2, 5
  31. 31.Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, San-jeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei. ImageNet Large Scale Visual Recognition Chal-lenge. International Journal of Computer Vision (IJCV), 115 (3):211–252, 2015. doi: 10.1007/s11263-015-0816-y. 4, 12
  32. 32.Peter Shaw, Jakob Uszkoreit, and Ashish Vaswani. Self-attention with relative position representations. arXiv preprint arXiv:1803.02155, 2018. 3
  33. 33.Kekai Sheng, Weiming Dong, Chongyang Ma, Xing Mei, Feiyue Huang, and Bao-Gang Hu. Attention-based multi-patch aggregation for image aesthetic assessment. In Pro-ceedings of the 26th ACM international conference on Mul-timedia, pages 879–886, 2018. 2, 6
  34. 34.Shaolin Su, Qingsen Yan, Yu Zhu, Cheng Zhang, Xin Ge, Jinqiu Sun, and Yanning Zhang. Blindly assess image qual-ity in the wild guided by a self-adaptive hyper network. In Proceedings of the IEEE/CVF Conference on Computer Vi-sion and Pattern Recognition, pages 3667–3676, 2020. 1, 2, 5, 6, 13
  35. 35.Chen Sun, Abhinav Shrivastava, Saurabh Singh, and Abhi-nav Gupta. Revisiting unreasonable effectiveness of data in deep learning era. In Proceedings of the IEEE International Conference on Computer Vision, pages 843–852, 2017. 12
  36. 36.Hossein Talebi and Peyman Milanfar. Nima: Neural image assessment. IEEE Transactions on Image Processing, 27(8): 3998–4011, 2018. 1, 2, 5, 6, 7
  37. 37.Bart Thomee, David A Shamma, Gerald Friedland, Ben-jamin Elizalde, Karl Ni, Douglas Poland, Damian Borth, and Li-Jia Li. Yfcc100m: The new data in multimedia research. Communications of the ACM, 59(2):64–73, 2016. 5
  38. 38.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszko-reit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Il-lia Polosukhin. Attention is all you need. In I. Guyon, U. V. Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vish-wanathan, and R. Garnett, editors, Advances in Neural Infor-mation Processing Systems, volume 30. Curran Associates, Inc., 2017. 1, 2, 4, 5, 11, 12
  39. 39.Jingtao Xu, Peng Ye, Qiaohong Li, Haiqing Du, Yong Liu, and David Doermann. Blind image quality assessment based on high order statistics aggregation. IEEE Transactions on Image Processing, 25(9):4444–4457, 2016. 6
  40. 40.Wufeng Xue, Lei Zhang, and Xuanqin Mou. Learning with-out human scores for blind image quality assessment. In Pro-ceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 995–1002, 2013. 2, 6
  41. 41.Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Russ R Salakhutdinov, and Quoc V Le. Xlnet: General-ized autoregressive pretraining for language understanding. In Advances in Neural Information Processing Systems 32, pages 5754–5764. Curran Associates, Inc., 2019. 2
  42. 42.Peng Ye, Jayant Kumar, Le Kang, and David Doermann. Un-supervised feature learning framework for no-reference im-age quality assessment. In Proceedings of the IEEE Con-ference on Computer Vision and Pattern Recognition, pages 1098–1105. IEEE, 2012. 2, 6
  43. 43.Zhenqiang Ying, Haoran Niu, Praful Gupta, Dhruv Maha-jan, Deepti Ghadiyaram, and Alan Bovik. From patches to pictures (paq-2-piq): Mapping the perceptual space of pic-ture quality. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3575– 3585, 2020. 1, 2, 5, 6
  44. 44.Hui Zeng, Lei Zhang, and Alan C Bovik. A probabilistic quality representation approach to deep blind image quality prediction. arXiv preprint arXiv:1708.08190, 2017. 6
  45. 45.Hui Zeng, Zisheng Cao, Lei Zhang, and Alan C Bovik. A unified probabilistic formulation of image aesthetic assess-ment. IEEE Transactions on Image Processing, 29:1548– 1561, 2019. 6
  46. 46.Hongyi Zhang, Moustapha Cisse, Yann N Dauphin, and David Lopez-Paz. mixup: Beyond empirical risk minimiza-tion. arXiv preprint arXiv:1710.09412, 2017. 5
  47. 47.Lin Zhang, Lei Zhang, and Alan C Bovik. A feature-enriched completely blind image quality evaluator. IEEE Transactions on Image Processing, 24(8):2579–2591, 2015. 2, 6
  48. 48.Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shecht-man, and Oliver Wang. The unreasonable effectiveness of deep features as a perceptual metric. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recogni-tion, pages 586–595, 2018. 2
  49. 49.Weixia Zhang, Kede Ma, Jia Yan, Dexiang Deng, and Zhou Wang. Blind image quality assessment using a deep bilinear convolutional neural network. IEEE Transactions on Cir-cuits and Systems for Video Technology, 30(1):36–47, 2018. 1, 2, 6
  50. 50.Hancheng Zhu, Leida Li, Jinjian Wu, Weisheng Dong, and Guangming Shi. Metaiqa: deep meta-learning for no-reference image quality assessment. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 14143–14152, 2020. 5, 6, 13

Citation

MLA
Ke, J., et al. “MUSIQ: Multi-scale Image Quality Transformer”. 2021 IEEE/CVF International Conference on Computer Vision (ICCV), 2021, pp. 5128–37, https://doi.org/10.1109/ICCV48922.2021.00510.
APA
Ke, J., Wang, Q., Wang, Y., Milanfar, P., & Yang, F. (2021). MUSIQ: Multi-scale Image Quality Transformer. 2021 IEEE/CVF International Conference on Computer Vision (ICCV), 5128–5137. https://doi.org/10.1109/ICCV48922.2021.00510
Chicago
Ke, J., Q. Wang, Y. Wang, P. Milanfar, and F. Yang. 2021. “MUSIQ: Multi-scale Image Quality Transformer”. 2021 IEEE/CVF International Conference on Computer Vision (ICCV), 5128–37. https://doi.org/10.1109/ICCV48922.2021.00510.
Harvard
Ke, J. et al. (2021) “MUSIQ: Multi-scale Image Quality Transformer”, 2021 IEEE/CVF International Conference on Computer Vision (ICCV). IEEE, pp. 5128–5137. Available at: https://doi.org/10.1109/ICCV48922.2021.00510.
Vancouver
1. Ke J, Wang Q, Wang Y, Milanfar P, Yang F (2021) MUSIQ: Multi-scale Image Quality Transformer. In: 2021 IEEE/CVF International Conference on Computer Vision (ICCV). IEEE, pp 5128–5137

BibTeX

@inproceedings{Ke_2021, title={MUSIQ: Multi-scale Image Quality Transformer}, url={http://dx.doi.org/10.1109/ICCV48922.2021.00510}, DOI={10.1109/iccv48922.2021.00510}, booktitle={2021 IEEE/CVF International Conference on Computer Vision (ICCV)}, publisher={IEEE}, author={Ke, Junjie and Wang, Qifei and Wang, Yilin and Milanfar, Peyman and Yang, Feng}, year={2021}, month=Oct, pages={5128–5137} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE