Coarse-To-Fine Deep Video Coding with Hyperprior-Guided Mode Prediction
Zhihao HuGuo LuJinyang GuoShan LiuWei JiangDong Xu
Presents a coarse-to-fine deep video compression framework that uses two-stage motion compensation alongside hyperprior-guided mode prediction to dynamically select block resolutions and skip residual coding without transmitting extra side information.
Rapid growth in video streaming and digital storage demands higher-efficiency compression systems to reduce network bandwidth and infrastructure costs. Although machine-learning-based video codecs have advanced significantly, existing models rely on single-scale motion compensation and struggle with complex motion scenarios. Additionally, deep learning approaches have struggled to efficiently adopt the adaptive mode-selection strategies used in traditional video standards without incurring severe computational overhead or transmitting expensive side information.
The article aims to design and evaluate an end-to-end deep video compression framework that produces superior motion compensation and compression efficiency without increasing bit rates or computational burdens. It introduces a two-stage coarse-to-fine framework coupled with hyperprior-guided adaptive mode prediction networks for both motion and residual data compression.
The researchers trained the complete architecture on the Vimeo-90K video dataset and evaluated performance across several standard industry benchmarks, including HEVC Class B through E, UVG, and MCL-JCV datasets. The approach uses a coarse level to capture broad motion patterns at one-fourth resolution, followed by a fine level that refines pixel-level predictions. Side statistical parameters (the mean and variance values already transmitted in the hyperprior stream) guide lightweight neural networks to adaptively select block coding resolutions and decide whether to skip residual transmission for flat or unchanged regions, requiring zero extra bits and negligible extra computation.
The evaluation yielded several key findings: First, the proposed framework consistently outperformed all competing learning-based video compression methods across standard datasets, exceeding recent models like ELF-VC by 0.5 dB in objective quality on the UVG dataset. Second, the framework achieved an average bit-rate saving of 4.58% compared to the traditional H.265/HEVC benchmark across the HEVC test sets. Third, on subjective quality metrics, the framework generally surpassed the newest international standard, VTM (Versatile Video Coding Test Model). Fourth, the system operated at 3.41 frames per second for high-definition 1080p video on a single graphics processing unit, making it over 3,000 times faster than the reference VTM software.
These results demonstrate that deep video coding can achieve commercial-grade compression efficiency while maintaining practical execution speeds. By eliminating the need for expensive rate-distortion search loops and dedicated mode-signaling bits, organizations can lower storage and data transmission expenses without compromising visual fidelity. The framework bridges the performance gap between traditional handcrafted standards and modern learned video systems.
Organizations developing next-generation video delivery architectures should consider integrating coarse-to-fine motion estimation and hyperprior-guided mode selection into their neural codec roadmaps. Further development should explore applying this hyperprior-guided mode prediction principle to other coding decisions and extending testing to broader video formats, real-time live streaming environments, and low-power edge hardware. While the current inference rate of 3.41 frames per second represents a major speedup over traditional reference software, additional optimization is needed before deployment in ultra-low-latency, real-time consumer streaming systems.
- Paper: Variational image compression with a scale hyperprior, Johannes Ballé et al. (2018). This foundational work introduced variational image compression with a scale hyperprior, establishing the hyperprior paradigm and latent statistical modeling that the source directly leverages for mode prediction.
- Paper: Joint Autoregressive and Hierarchical Priors for Learned Image Compression, David Minnen et al. (2018). This paper establishes the joint autoregressive and hierarchical hyperprior framework that underpins the mean and variance estimation used for entropy modeling in neural compression.
- Paper: End-to-end Optimized Image Compression, Johannes Ballé et al. (2016). This paper provides the foundational principles of end-to-end optimized learned transform coding and continuous relaxations for quantization upon which deep video codecs build.
- Paper: Hierarchical Model-Based Motion Estimation, J. Bergen et al. (1992). This classic work details the coarse-to-fine multiresolution hierarchy for model-based motion estimation and compensation that motivates the source's coarse-to-fine deep video framework.
- Paper: Optical Flow Estimation Using a Spatial Pyramid Network, Anurag Ranjan et al. (2016). This paper develops coarse-to-fine spatial pyramid warping for deep motion estimation, providing essential context for multi-scale optical flow alignment in neural video coding.
- Paper: Learned Image Compression With Discretized Gaussian Mixture Likelihoods and Attention Modules, Zhengxue Cheng et al. (2020). This work demonstrates advanced learned latent likelihood modeling and attention modules for rate estimation, reinforcing the statistical coding concepts adapted in the source.
- Paper: C3: High-Performance and Low-Complexity Neural Compression from a Single Image or Video, Hyunjik Kim et al. (2024). This paper explores an alternative paradigm in neural media coding by optimizing lightweight instance-adaptive models to minimize decoding complexity in learned video compression.
