C3: High-Performance and Low-Complexity Neural Compression from a Single Image or Video
Hyunjik KimMatthias BauerLucas TheisJonathan Richard SchwarzEmilien Dupont
Presents an instance-overfitted neural image and video compression method that matches state-of-the-art codec quality while drastically reducing decoding complexity to under 5k multiply-accumulate operations per pixel.
Neural network-based data compression offers strong compression efficiency but often demands massive computational power during decoding. This high computational burden creates a severe bottleneck on resource-constrained devices, such as smartphones, making real-time playback and broad adoption challenging. To solve this dilemma, the article introduces C3, an approach that optimizes very small neural networks tailored to individual images or videos rather than relying on a single, massive model trained across large datasets.
The article evaluates whether overfitting small neural models directly to individual media files can match state-of-the-art compression efficiency while slashing decoding complexity. The researchers designed a unified framework for images and videos that builds on earlier instance-based architectures by incorporating smooth quantization approximations, tailored noise distributions, refined network layers, and spatial-temporal video patch modeling. They benchmarked this approach against industry-standard classical formats and top-tier neural codecs across standard image and video evaluation datasets.
The analysis reveals four key findings. First, on the CLIC2020 image benchmark, C3 matches the compression efficiency of the reference next-generation codec VTM while requiring less than 3,000 multiply-accumulate operations per pixel—an order of magnitude lower than comparable neural decoders. Second, tailoring architecture choices to specific image instances reduces the required bitrate relative to VTM by approximately 2.9%. Third, extending the method to video on the UVG benchmark matches the compression quality of the Video Compression Transformer while utilizing only 4,400 operations per pixel, representing less than 0.1% of the baseline neural decoding cost. Finally, ablation experiments confirm that smooth quantization rounding, Kumaraswamy noise injection, and expressive activation functions provide the vast majority of these performance improvements.
These findings demonstrate that neural video and image compression can achieve top-tier compression efficiency without imposing heavy decoding hardware requirements. This breakthrough substantially lowers playback costs and power consumption, making neural codecs viable for low-power streaming devices. While instance-tailored compression requires significant encoding time up front, it is especially valuable for asymmetric distribution workflows, such as on-demand streaming platforms where content is compressed once but played back millions of times.
Organizations evaluating neural media pipelines should consider instance-overfitted architectures for high-volume streaming distributions where low-power client decoding is paramount. Before wide deployment, development teams must address the substantial encoding overhead by testing faster optimization schedules and exploring parallel decoding techniques. Because evaluations were conducted on standard research datasets using unoptimized code, further validation on production video streams and diverse hardware architectures is recommended.
- Paper: End-to-end Optimized Image Compression, Johannes Ballé et al. (2016). This establishes the end-to-end learned compression framework, including additive-noise quantization, that C3 adapts while pursuing far lower decoding complexity.
- Paper: Variational image compression with a scale hyperprior, Johannes Ballé et al. (2018). Its learned scale hyperprior provides the entropy-modeling foundation for understanding how C3 estimates compact latent representations and side information.
- Paper: Joint Autoregressive and Hierarchical Priors for Learned Image Compression, David Minnen et al. (2018). The joint autoregressive and hierarchical prior architecture supplies the probability-modeling advances against which C3’s simpler, instance-tailored decoding strategy is motivated.
- Paper: ELIC: Efficient Learned Image Compression with Unevenly Grouped Space-Channel Contextual Adaptive Coding, Dailan He et al. (2022). ELIC makes the learned-compression quality–latency trade-off explicit, clarifying the practical decoding bottleneck that C3 is designed to overcome.
- Paper: Learned Image Compression With Discretized Gaussian Mixture Likelihoods and Attention Modules, Zhengxue Cheng et al. (2020). Its discretized likelihoods and attention-based autoencoder illustrate the high-efficiency image-codec design that C3 simplifies through per-instance optimization.
No sufficiently relevant recommendations were found.
