Coarse-To-Fine Deep Video Coding with Hyperprior-Guided Mode Prediction

Zhihao HuGuo LuJinyang GuoShan LiuWei JiangDong Xu

article2022CVPR128 citations

Presents a coarse-to-fine deep video compression framework that uses two-stage motion compensation alongside hyperprior-guided mode prediction to dynamically select block resolutions and skip residual coding without transmitting extra side information.

Listen

Rapid growth in video streaming and digital storage demands higher-efficiency compression systems to reduce network bandwidth and infrastructure costs. Although machine-learning-based video codecs have advanced significantly, existing models rely on single-scale motion compensation and struggle with complex motion scenarios. Additionally, deep learning approaches have struggled to efficiently adopt the adaptive mode-selection strategies used in traditional video standards without incurring severe computational overhead or transmitting expensive side information.

The article aims to design and evaluate an end-to-end deep video compression framework that produces superior motion compensation and compression efficiency without increasing bit rates or computational burdens. It introduces a two-stage coarse-to-fine framework coupled with hyperprior-guided adaptive mode prediction networks for both motion and residual data compression.

The researchers trained the complete architecture on the Vimeo-90K video dataset and evaluated performance across several standard industry benchmarks, including HEVC Class B through E, UVG, and MCL-JCV datasets. The approach uses a coarse level to capture broad motion patterns at one-fourth resolution, followed by a fine level that refines pixel-level predictions. Side statistical parameters (the mean and variance values already transmitted in the hyperprior stream) guide lightweight neural networks to adaptively select block coding resolutions and decide whether to skip residual transmission for flat or unchanged regions, requiring zero extra bits and negligible extra computation.

The evaluation yielded several key findings: First, the proposed framework consistently outperformed all competing learning-based video compression methods across standard datasets, exceeding recent models like ELF-VC by 0.5 dB in objective quality on the UVG dataset. Second, the framework achieved an average bit-rate saving of 4.58% compared to the traditional H.265/HEVC benchmark across the HEVC test sets. Third, on subjective quality metrics, the framework generally surpassed the newest international standard, VTM (Versatile Video Coding Test Model). Fourth, the system operated at 3.41 frames per second for high-definition 1080p video on a single graphics processing unit, making it over 3,000 times faster than the reference VTM software.

These results demonstrate that deep video coding can achieve commercial-grade compression efficiency while maintaining practical execution speeds. By eliminating the need for expensive rate-distortion search loops and dedicated mode-signaling bits, organizations can lower storage and data transmission expenses without compromising visual fidelity. The framework bridges the performance gap between traditional handcrafted standards and modern learned video systems.

Organizations developing next-generation video delivery architectures should consider integrating coarse-to-fine motion estimation and hyperprior-guided mode selection into their neural codec roadmaps. Further development should explore applying this hyperprior-guided mode prediction principle to other coding decisions and extending testing to broader video formats, real-time live streaming environments, and low-power edge hardware. While the current inference rate of 3.41 frames per second represents a major speedup over traditional reference software, additional optimization is needed before deployment in ultra-low-latency, real-time consumer streaming systems.

arXiv: 2206.07460
Cover for Coarse-To-Fine Deep Video Coding with Hyperprior-Guided Mode Prediction

Abstract

The previous deep video compression approaches only use the single scale motion compensation strategy and rarely adopt the mode prediction technique from the traditional standards like H.264/H.265 for both motion and residual compression. In this work, we first propose a coarse-to-fine (C2F) deep video compression framework for better motion compensation, in which we perform motion estimation, compression and compensation twice in a coarse to fine manner. Our C2F framework can achieve better motion compensation results without significantly increasing bit costs. Observing hyperprior information (i.e., the mean and variance values) from the hyperprior networks contains discriminant statistical information of different patches, we also propose two efficient hyperprior-guided mode prediction methods. Specifically, using hyperprior information as the input, we propose two mode prediction networks to respectively predict the optimal block resolutions for better motion coding and decide whether to skip residual information from each block for better residual coding without introducing additional bit cost while bringing negligible extra computation cost. Comprehensive experimental results demonstrate our proposed C2F video compression framework equipped with the new hyperprior-guided mode prediction methods achieves the state-of-the-art performance on HEVC, UVG and MCL-JCV datasets.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 2.1. Image and Video Compression
  • 2.2. Coarse-to-fine Strategy in Computer Vision
  • 3. Method
  • 3.1. Overview
  • 3.2. Coarse-to-Fine Motion Compensation
  • 3.3. Hyperprior-guided Adaptive Motion Compression and Residual Compression
  • 3.4. Loss Function and Entropy Coding
  • 4. Experiments
  • 4.1. Experimental Setup
  • 4.2. Experimental Results
  • 4.3. Ablation Study
  • 4.4. Model Analysis
  • 5. Conclusion
  • References

Knowls

  1. Knowl 1 — Coarse-to-fine feature-space video compression framework

    model/method

    The proposed video codec compresses each P-frame in feature space using two-stage coarse-to-fine motion compensation followed by residual compression. An input frame XtX_t and the previously reconstructed reference frame X^t−1\hat{X}_{t-1} are transformed into an input feature FtF_t and a reference feature Ft−1refF^{\mathrm{ref}}_{t-1} by a feature-extraction network.

    In the coarse branch, both features are downsampled by a factor of four. Motion estimation, motion compression, and upsampling produce a reconstructed coarse offset map, which controls a deformable convolution that warps Ft−1refF^{\mathrm{ref}}_{t-1}. The warped feature is concatenated with the reference feature and processed by convolutions to produce an intermediate predicted feature F~t\tilde{F}_t.

    The fine branch uses F~t\tilde{F}_t as its reference and FtF_t as its target. It performs motion estimation, motion compression, and deformable-convolution-based motion compensation again at the original feature resolution, producing the final predicted feature Fˉt\bar{F}_t. The residual feature is Rt=Ft−FˉtR_t=F_t-\bar{F}_t and is compressed separately. The reconstructed feature F^t=Fˉt+R^t\hat{F}_t=\bar{F}_t+\hat{R}_t is passed through a frame-reconstruction network to produce X^t\hat{X}_t, which is stored for coding the next frame. The coarse branch adds only a low-resolution motion representation, while the fine branch provides detailed motion refinement.

  2. Knowl 2 — Hyperprior-guided resolution-mode prediction for motion coding

    model/method

    The hyperprior-guided adaptive motion compression (HAMC) method predicts a spatially and channel-adaptive resolution for coding the fine-level encoded motion feature MtM_t. Its mode-prediction network receives the mean and variance decoded by the motion hyperprior network; these statistics describe the local distribution of the motion feature and are already available at both encoder and decoder, so the predicted modes require no separately transmitted mode map.

    For an encoded feature with 128 channels, the predictor produces four confidence scores per channel for each candidate mode. It has two prediction branches: one predicts a mode for each 2×22\times2 subblock and the other predicts a mode for the enclosing block, such as a 4×44\times4 block. Four basic resolution modes are available at each block scale. During training, differentiable Gumbel-softmax selection is used; during inference, the mode with maximum confidence is selected.

    If the enclosing block selects the mode that allows subblock refinement, the four independently predicted 2×22\times2 modes are combined with the enclosing-block prediction to form the final partition. Otherwise, the enclosing-block mode is used directly. This lets the codec use fine partitions where motion statistics vary strongly and coarse partitions where motion is smooth.

  3. Knowl 3 — Mode-guided motion-feature quantization and transmission

    algorithm

    HAMC transforms the motion-estimation offset map OtO_t into an encoded motion feature MtM_t, then uses the hyperprior-guided resolution mode to reduce the number of transmitted motion values.

    Input: offset map OtO_t, decoded hyperprior mean and variance, and the mode-prediction network.

    Output: quantized reconstructed motion feature M^t\hat{M}_t and its arithmetic-coded bitstream.

    Input: offset map O_t and hyperprior mean/variance
    Encode O_t with the motion encoder to obtain M_t
    Predict a resolution mode for each block and channel
    For each selected block:
        If the mode groups multiple values, average-pool those values into one value
        Quantize the resulting value
        Arithmetic-code the quantized value
    Transmit the arithmetic-coded quantized values
    At the decoder, arithmetic-decode each transmitted value
    For each grouped block, upsample the decoded value to all positions in that block
    Return the reconstructed motion feature M_hat_t

    For example, four values in a 2×22\times2 block can be replaced by their average before quantization and entropy coding, and the decoded average is replicated over the four positions. Larger blocks therefore save bits in smooth regions, whereas smaller blocks preserve motion detail near object boundaries. The predicted modes are inferred from hyperprior information and introduce no additional bit cost.

  4. Knowl 4 — Hyperprior-guided adaptive residual compression

    model/method

    The hyperprior-guided adaptive residual compression (HARC) method exploits the sparsity of the residual feature Rt=Ft−FˉtR_t=F_t-\bar{F}_t. A one-branch mode-prediction network takes the decoded mean and variance from the residual hyperprior network and predicts a binary “skip” or “non-skip” mode for each residual-feature entry and channel.

    Entries or blocks with statistically insignificant residual information are assigned the skip mode and are not transmitted. Entries with significant residual information are assigned the non-skip mode and are quantized and entropy-coded. Because the mode predictor uses hyperprior statistics that are already decoded by both sides, HARC does not transmit an additional mode map and adds negligible computation. It reduces residual bit cost while retaining information needed for reconstructing regions with substantial residual structure.

  5. Knowl 5 — End-to-end rate-distortion objective

    equation

    The codec is trained end-to-end with the rate-distortion objective

    L=H(M^tc)+H(M^t)+H(Y^t)+λd(Xt,X^t).\mathcal{L}=H(\hat{M}^{c}_t)+H(\hat{M}_t)+H(\hat{Y}_t)+\lambda d(X_t,\hat{X}_t).

    Here tt indexes the video frame, M^tc\hat{M}^{c}_t is the quantized coarse-level motion feature, M^t\hat{M}_t is the quantized fine-level motion feature, and Y^t\hat{Y}_t is the quantized encoded residual feature. The function H(⋅)H(\cdot) estimates the number of bits required to entropy-code its argument, XtX_t is the input frame, X^t\hat{X}_t is the reconstructed frame, d(⋅,⋅)d(\cdot,\cdot) is the frame-distortion measure, and λ>0\lambda>0 controls the rate-distortion trade-off.

    During training, the fine-motion and residual bit rates are estimated with a hyperprior-based bit-rate network without a time-consuming autoregressive model. The coarse-motion rate uses a simpler bit-rate estimator because the coarse motion feature has relatively small spatial resolution.

  6. Knowl 6 — Training and evaluation protocol

    experimental setup

    The model is trained on Vimeo-90K, which contains 89,800 seven-frame sequences at 488×256488\times256 resolution. Sequences are randomly flipped and cropped to 256×256256\times256 patches.

    Training uses three stages: (1) 2,000,000 steps with two consecutive frames, one I-frame and one P-frame, without HAMC or HARC; (2) 300,000 additional steps with seven-frame sequences; and (3) 200,000 steps after adding both HAMC and HARC. The initial learning rate is 5×10−55\times10^{-5} and is reduced by 80% at steps 1,900,000 and 2,400,000. Adam is used with batch size 4 in the first stage and 2 thereafter. Mean squared error is used for PSNR models, while MS-SSIM models receive an additional 100,000-step fine-tuning phase using MS-SSIM distortion.

    Evaluation uses HEVC Classes B, C, D, and E, UVG, and MCL-JCV. PSNR, MS-SSIM, and bits per pixel (bpp) are reported. Comparisons use low-delay HM and VTM configurations with GoP size 100 and BPG-compressed I-frames; a stronger P-frame model is used for every fourth frame to reduce error accumulation.

  7. Knowl 7 — Compression performance against learned and conventional codecs

    empirical result

    Across the UVG, MCL-JCV, and HEVC Class B, C, D, and E datasets, the complete coarse-to-fine codec with HAMC and HARC outperforms the compared learning-based codecs by a large margin in PSNR. On UVG, it improves over ELF-VC by approximately 0.50.5 dB at 0.10.1 bpp. It achieves better PSNR than H.265(HM) on most datasets and is close to VTM at high bit rates on the high-resolution UVG, MCL-JCV, HEVC Class B, and HEVC Class E datasets.

    For MS-SSIM, the complete model generally outperforms the compared learning-based methods and is reported to outperform VTM in the overall comparison. Using H.265(HM) as the anchor, the average bit-rate saving over HEVC Classes B, C, and D is 4.58%4.58\%. These results support the claim that coarse-to-fine compensation improves reconstruction quality without requiring a proportionate increase in motion bits.

  8. Knowl 8 — Ablation of coarse-to-fine compensation and adaptive coding

    empirical result

    On HEVC Class E, the re-implemented FVC codec is used as the baseline. Adding only the coarse-to-fine motion-compensation framework improves PSNR by 0.30.3 dB at 0.080.08 bpp. Adding HAMC to the coarse-to-fine framework yields a further 0.60.6 dB improvement at 0.030.03 bpp. Adding HARC as well produces the best result, outperforming the baseline by 1.21.2 dB at 0.030.03 bpp.

    The coarse-to-fine framework is more beneficial at higher bit rates because it improves the quality of the predicted feature, whereas HAMC provides larger gains at lower bit rates because it removes motion-coding redundancy. The final improvement demonstrates that the motion-compensation architecture, adaptive motion coding, and adaptive residual coding contribute complementary gains.

  9. Knowl 9 — Motion-compensation quality from the two-stage representation

    empirical result

    A motion visualization on a reconstructed P-frame from HEVC Class B compares the single-scale FVC baseline with the proposed two-stage motion representation. The coarse-level offset map captures patch-level motion and costs 0.0030.003 bpp, while the fine-level offset map captures more detailed, pixel-level motion and costs 0.0140.014 bpp. The single-scale baseline offset map costs 0.0170.017 bpp.

    Despite a similar total motion bit cost, the coarse-to-fine representation improves the motion-compensation PSNR from 28.828.8 dB for the baseline to 30.530.5 dB, a gain of 1.71.7 dB. On HEVC Class E, the average motion-compensation improvement is 1.141.14 dB at similar bpp. The results indicate that a low-cost coarse alignment allows the fine stage to spend its capacity on residual motion details and handle complex motion patterns more accurately.

  10. Knowl 10 — Negligible computational overhead of hyperprior-guided modes

    empirical result

    On 1920×10801920\times1080 videos processed with a single NVIDIA 2080Ti GPU, the complete C2F+HAMC+HARC codec runs at 3.413.41 frames per second, while the C2F codec without HAMC or HARC runs at 3.433.43 frames per second. Thus, the two hyperprior-guided mode-prediction networks reduce coding cost without a measurable practical loss in inference speed. The paper also reports that the complete method is about 3,000 times faster than VTM, whose speed is below 0.0010.001 frames per second under the comparison configuration.

Coverage note — No substantial contributed method or result was deliberately omitted; implementation-level details of inherited feature-extraction, frame-reconstruction, and entropy networks were excluded because they are reused components rather than contributions of this paper.

References

  1. 1.HEVC test model (HM). https : / / hevc . hhi . fraunhofer.de/HM-doc/. Accessed: 2022-03-28. 2, 7
  2. 2.VVC test model (VTM). https : / / jvet . hhi . fraunhofer.de/. Accessed: 2022-03-28. 2, 7
  3. 3.Eirikur Agustsson, David Minnen, Nick Johnston, Johannes Balle, Sung Jin Hwang, and George Toderici. Scale-space flow for end-to-end optimized video compression. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 8503–8512, 2020. 1, 2, 7
  4. 4.Johannes Balle, Valero Laparra, and Eero P Simoncelli. End-to-end optimized image compression. International Conference on Learning Representations (ICLR), 2017. 2, 6
  5. 5.Johannes Balle, David Minnen, Saurabh Singh, Sung Jin Hwang, and Nick Johnston. Variational image compression with a scale hyperprior. International Conference on Learning Representations (ICLR), 2018. 2
  6. 6.Fabrice Bellard. BPG image format. URL https://bellard.org/bpg, 2015. 2, 7
  7. 7.Zhenghao Chen, Shuhang Gu, Guo Lu, and Dong Xu. Exploiting intra-slice and inter-slice redundancy for learning-based lossless volumetric image compression. IEEE Transactions on Image Processing, 2022. 2
  8. 8.Zhenghao Chen, Guo Lu, Zhihao Hu, Shan Liu, Wei Jiang, and Dong Xu. LSVC: A learning-based stereo video compression framework. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 734–743, 2022. 2
  9. 9.Zhengxue Cheng, Heming Sun, Masaru Takeuchi, and Jiro Katto. Learning image and video compression through spatial-temporal energy compaction. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 10071–10080, 2019. 2
  10. 10.Zhengxue Cheng, Heming Sun, Masaru Takeuchi, and Jiro Katto. Learned image compression with discretized gaussian mixture likelihoods and attention modules. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 7939–7948, 2020. 2
  11. 11.Jifeng Dai, Haozhi Qi, Y. Xiong, Y. Li, Guodong Zhang, H. Hu, and Y. Wei. Deformable convolutional networks. 2017 IEEE International Conference on Computer Vision (ICCV), pages 764–773, 2017. 4
  12. 12.Abdelaziz Djelouah, Joaquim Campos, Simone Schaub-Meyer, and Christopher Schroers. Neural inter-frame compression for video coding. In Proceedings of the IEEE International Conference on Computer Vision, pages 6421–6429, 2019. 1, 2
  13. 13.Runsen Feng, Yaojun Wu, Zongyu Guo, Zhizheng Zhang, and Zhibo Chen. Learned video compression with feature-level residuals. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, pages 120–121, 2020. 2
  14. 14.Zhihao Hu, Zhenghao Chen, Dong Xu, Guo Lu, Wanli Ouyang, and Shuhang Gu. Improving deep video compression by resolution-adaptive flow coding. In European Conference on Computer Vision, pages 193–209. Springer, 2020. 2, 7
  15. 15.Zhihao Hu, Guo Lu, and Dong Xu. FVC: A new framework towards deep video compression in feature space. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1502–1511, 2021. 1, 2, 3, 4, 6, 7
  16. 16.Tak-Wai Hui, Xiaoou Tang, and Chen Change Loy. Lite-flownet: A lightweight convolutional neural network for optical flow estimation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 8981–8989, 2018. 2
  17. 17.Eric Jang, Shixiang Gu, and Ben Poole. Categorical reparameterization with gumbel-softmax. International Conference on Learning Representations (ICLR), 2017. 5
  18. 18.Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization. International Conference for Learning Representations, 2015. 6
  19. 19.Jiahao Li, Bin Li, and Yan Lu. Deep contextual video compression. Advances in Neural Information Processing Systems, 34, 2021. 7
  20. 20.Jianping Lin, Dong Liu, Houqiang Li, and Feng Wu. M-LVC: Multiple frames prediction for learned video compression. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 3546–3554, 2020. 2
  21. 21.Salvator Lombardo, Jun Han, Christopher Schroers, and Stephan Mandt. Deep generative video compression. In Advances in Neural Information Processing Systems, pages 9287–9298, 2019. 2
  22. 22.Guo Lu, Chunlei Cai, Xiaoyun Zhang, Li Chen, Wanli Ouyang, Dong Xu, and Zhiyong Gao. Content adaptive and error propagation aware deep video compression. In European Conference on Computer Vision, pages 456–472. Springer, 2020. 2
  23. 23.Guo Lu, Wanli Ouyang, Dong Xu, Xiaoyun Zhang, Chunlei Cai, and Zhiyong Gao. DVC: An end-to-end deep video compression framework. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 11006–11015, 2019. 1, 2, 6
  24. 24.Guo Lu, Xiaoyun Zhang, Wanli Ouyang, Li Chen, Zhiyong Gao, and Dong Xu. An end-to-end learning framework for video compression. IEEE Transactions on Pattern Analysis and Machine Intelligence, in Press:1–1, 2020. 2
  25. 25.Fabian Mentzer, Eirikur Agustsson, Johannes Balle, David Minnen, Nick Johnston, and George Toderici. Towards generative video compression. arXiv preprint arXiv:2107.12038, 2021. 2, 7
  26. 26.A. Mercat, Marko Viitanen, and J. Vanne. UVG dataset: 50/120fps 4k sequences for video codec analysis and development. Proceedings of the 11th ACM Multimedia Systems Conference, 2020. 6
  27. 27.David Minnen, Johannes Balle, and George D Toderici. Joint autoregressive and hierarchical priors for learned image compression. In Advances in Neural Information Processing Systems, pages 10771–10780, 2018. 2, 5
  28. 28.David Minnen and Saurabh Singh. Channel-wise autoregressive entropy models for learned image compression. In 2020 IEEE International Conference on Image Processing (ICIP), pages 3339–3343. IEEE, 2020. 2
  29. 29.Anurag Ranjan and Michael J Black. Optical flow estimation using a spatial pyramid network. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 4161–4170, 2017. 1, 2
  30. 30.Oren Rippel, Alexander G Anderson, Kedar Tatwawadi, Sanjay Nair, Craig Lytle, and Lubomir Bourdev. ELF-VC: Efficient learned flexible-rate video coding. In Proceedings of the IEEE International Conference on Computer Vision, pages 3033–3042, 2021. 2, 7
  31. 31.Oren Rippel, Sanjay Nair, Carissa Lew, Steve Branson, Alexander G Anderson, and Lubomir Bourdev. Learned video compression. In Proceedings of the IEEE International Conference on Computer Vision, pages 3454–3463, 2019. 2
  32. 32.Gary Sullivan. Versatile video coding (VVC) arrives. In 2020 IEEE International Conference on Visual Communications and Image Processing (VCIP), pages 1–1. IEEE, 2020. 1, 2
  33. 33.Gary J Sullivan, Jens-Rainer Ohm, Woo-Jin Han, and Thomas Wiegand. Overview of the high efficiency video coding (HEVC) standard. IEEE Transactions on circuits and systems for video technology, 22(12):1649–1668, 2012. 1, 2, 6
  34. 34.Deqing Sun, Xiaodong Yang, Ming-Yu Liu, and Jan Kautz. Pwc-net: Cnns for optical flow using pyramid, warping, and cost volume. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 8934–8943, 2018. 2
  35. 35.David S Taubman and Michael W Marcellin. Jpeg2000: Standard for interactive imaging. Proceedings of the IEEE, 90(8):1336–1357, 2002. 2
  36. 36.Lucas Theis, Wenzhe Shi, Andrew Cunningham, and Ferenc Huszar. Lossy image compression with compressive autoencoders. International Conference for Learning Representations, 2017. 2
  37. 37.George Toderici, Sean M O’Malley, Sung Jin Hwang, Damien Vincent, David Minnen, Shumeet Baluja, Michele Covell, and Rahul Sukthankar. Variable rate image compression with recurrent neural networks. International Conference for Learning Representations, 2017. 2
  38. 38.George Toderici, Damien Vincent, Nick Johnston, Sung Jin Hwang, David Minnen, Joel Shor, and Michele Covell. Full resolution image compression with recurrent neural networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 5306–5314, 2017. 2
  39. 39.Gregory K Wallace. The JPEG still picture compression standard. IEEE transactions on consumer electronics, 38(1):xviii–xxxiv, 1992. 2
  40. 40.Haiqiang Wang, Weihao Gan, Sudeng Hu, Joe Yuchieh Lin, Lina Jin, Longguang Song, Ping Wang, Ioannis Katsavounidis, Anne Aaron, and C-C Jay Kuo. MCL-JCV: a jnd-based H.264/AVC video quality assessment dataset. In 2016 IEEE International Conference on Image Processing (ICIP), pages 1509–1513. IEEE, 2016. 6
  41. 41.Xintao Wang, Kelvin C. K. Chan, K. Yu, C. Dong, and Chen Change Loy. EDVR: Video restoration with enhanced deformable convolutional networks. 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), pages 1954–1963, 2019. 1, 2
  42. 42.Zhou Wang, Eero P Simoncelli, and Alan C Bovik. Multi-scale structural similarity for image quality assessment. In The Thrity-Seventh Asilomar Conference on Signals, Systems & Computers, 2003, volume 2, pages 1398–1402. Ieee, 2003. 6
  43. 43.Thomas Wiegand, Gary J Sullivan, Gisle Bjontegaard, and Ajay Luthra. Overview of the H.264/AVC video coding standard. IEEE Transactions on circuits and systems for video technology, 13(7):560–576, 2003. 1, 2
  44. 44.Chao-Yuan Wu, Nayan Singhal, and Philipp Krahenbuhl. Video compression through image interpolation. In Proceedings of the European Conference on Computer Vision (ECCV), pages 416–431, 2018. 2
  45. 45.Tianfan Xue, Baian Chen, Jiajun Wu, Donglai Wei, and William T Freeman. Video enhancement with task-oriented flow. International Journal of Computer Vision, 127(8):1106–1125, 2019. 6
  46. 46.Ren Yang, Fabian Mentzer, Luc Van Gool, and Radu Timofte. Learning for video compression with hierarchical quality and recurrent enhancement. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6628–6637, 2020. 2
  47. 47.R. Yang, Fabian Mentzer, L. Van Gool, and R. Timofte. Learning for video compression with recurrent auto-encoder and recurrent probability model. IEEE Journal of Selected Topics in Signal Processing, 15:388–401, 2021. 7

Citation

MLA
Hu, Z., et al. “Coarse-to-fine Deep Video Coding with Hyperprior-guided Mode Prediction”. arXiv, 2022, http://arxiv.org/abs/2206.07460v1.
APA
Hu, Z., Lu, G., Guo, J., Liu, S., Jiang, W., & Xu, D. (2022). Coarse-to-fine Deep Video Coding with Hyperprior-guided Mode Prediction. arXiv. http://arxiv.org/abs/2206.07460v1
Chicago
Hu, Z., G. Lu, J. Guo, S. Liu, W. Jiang, and D. Xu. 2022. “Coarse-to-fine Deep Video Coding with Hyperprior-guided Mode Prediction”. arXiv. http://arxiv.org/abs/2206.07460v1.
Harvard
Hu, Z. et al. (2022) “Coarse-to-fine Deep Video Coding with Hyperprior-guided Mode Prediction”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2206.07460v1.
Vancouver
1. Hu Z, Lu G, Guo J, Liu S, Jiang W, Xu D (2022) Coarse-to-fine Deep Video Coding with Hyperprior-guided Mode Prediction. arXiv

BibTeX

@article{hu2022coarse,
  title = {Coarse-to-fine Deep Video Coding with Hyperprior-guided Mode Prediction},
  author = {Hu, Zhihao and Lu, Guo and Guo, Jinyang and Liu, Shan and Jiang, Wei and Xu, Dong},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2206.07460v1},
  eprint = {2206.07460}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE