C3: High-Performance and Low-Complexity Neural Compression from a Single Image or Video

Hyunjik KimMatthias BauerLucas TheisJonathan Richard SchwarzEmilien Dupont

article2024CVPR100 citations

Presents an instance-overfitted neural image and video compression method that matches state-of-the-art codec quality while drastically reducing decoding complexity to under 5k multiply-accumulate operations per pixel.

Listen

Neural network-based data compression offers strong compression efficiency but often demands massive computational power during decoding. This high computational burden creates a severe bottleneck on resource-constrained devices, such as smartphones, making real-time playback and broad adoption challenging. To solve this dilemma, the article introduces C3, an approach that optimizes very small neural networks tailored to individual images or videos rather than relying on a single, massive model trained across large datasets.

The article evaluates whether overfitting small neural models directly to individual media files can match state-of-the-art compression efficiency while slashing decoding complexity. The researchers designed a unified framework for images and videos that builds on earlier instance-based architectures by incorporating smooth quantization approximations, tailored noise distributions, refined network layers, and spatial-temporal video patch modeling. They benchmarked this approach against industry-standard classical formats and top-tier neural codecs across standard image and video evaluation datasets.

The analysis reveals four key findings. First, on the CLIC2020 image benchmark, C3 matches the compression efficiency of the reference next-generation codec VTM while requiring less than 3,000 multiply-accumulate operations per pixel—an order of magnitude lower than comparable neural decoders. Second, tailoring architecture choices to specific image instances reduces the required bitrate relative to VTM by approximately 2.9%. Third, extending the method to video on the UVG benchmark matches the compression quality of the Video Compression Transformer while utilizing only 4,400 operations per pixel, representing less than 0.1% of the baseline neural decoding cost. Finally, ablation experiments confirm that smooth quantization rounding, Kumaraswamy noise injection, and expressive activation functions provide the vast majority of these performance improvements.

These findings demonstrate that neural video and image compression can achieve top-tier compression efficiency without imposing heavy decoding hardware requirements. This breakthrough substantially lowers playback costs and power consumption, making neural codecs viable for low-power streaming devices. While instance-tailored compression requires significant encoding time up front, it is especially valuable for asymmetric distribution workflows, such as on-demand streaming platforms where content is compressed once but played back millions of times.

Organizations evaluating neural media pipelines should consider instance-overfitted architectures for high-volume streaming distributions where low-power client decoding is paramount. Before wide deployment, development teams must address the substantial encoding overhead by testing faster optimization schedules and exploring parallel decoding techniques. Because evaluations were conducted on standard research datasets using unoptimized code, further validation on production video streams and diverse hardware architectures is recommended.

No sufficiently relevant recommendations were found.

Cover for C3: High-Performance and Low-Complexity Neural Compression from a Single Image or Video

Abstract

Most neural compression models are trained on large datasets of images or videos in order to generalize to unseen data. Such generalization typically requires large and expressive architectures with a high decoding complexity. Here we introduce C3, a neural compression method with strong rate-distortion (RD) performance that instead overfits a small model to each image or video separately. The resulting decoding complexity of C3 can be an order of magnitude lower than neural baselines with similar RD performance. C3 builds on Cool-chic [43] and makes several simple and effective improvements for images. We further develop new methodology to apply C3 to videos. On the CLIC2020 image benchmark, we match the RD performance of VTM, the reference implementation of the H.266 codec, with less than 3k MACs/pixel for decoding. On the UVG video benchmark, we match the RD performance of the Video Compression Transformer [60], a well-established neural video codec, with less than 5k MACs/pixel for decoding.

Table of Contents

  • 1. Introduction
  • 2. Background: Cool-chic
  • 3. C3: Improving Cool-chic
  • 3.1. Optimization improvements
  • 3.2. Model improvements
  • 3.3. Video-specific methodology
  • 4. Related work
  • 5. Results
  • 5.1. Image compression
  • 6. Conclusion, limitations and future work
  • References

Knowls

  1. Knowl 1 — C3 Architecture and Per-Instance Rate-Distortion Optimization Objective

    model/method

    C3 (Cooler-ChiC) compresses an individual image or video instance xx by overfitting a hierarchy of continuous latent grids and lightweight neural networks directly to that single instance via gradient descent, bypassing amortized encoder transforms.

    The image architecture comprises three main components:

    1. Multi-resolution latent grids z=(z1,…,zN)z = (z^1, \dots, z^N), structured at spatial resolutions (h,w),(h/2,w/2),…,(h/2N−1,w/2N−1)(h, w), (h/2, w/2), \dots, (h/2^{N-1}, w/2^{N-1}) for an image of dimensions h×wh \times w.
    2. A synthesis network fθf_\theta (a small convolutional network of depth ≤4\le 4 and width ≤40\le 40) that maps the concatenated, deterministically or learned upsampled latent grids Up(z)∈Rh×w×N\text{Up}(z) \in \mathbb{R}^{h \times w \times N} to the reconstructed image xrec=fθ(Up(z))x_{\text{rec}} = f_\theta(\text{Up}(z)).
    3. An autoregressive entropy network gψg_\psi that predicts the location and scale parameters (μijn,σijn)(\mu^n_{ij}, \sigma^n_{ij}) of an integrated Laplace distribution for each latent element zijnz^n_{ij} using a causally masked local spatial neighborhood:

    Pψ(zn)=∏i,jP(zijn;μijn,σijn)P_\psi(z^n) = \prod_{i,j} P(z^n_{ij}; \mu^n_{ij}, \sigma^n_{ij})

    μijn,σijn=gψ(context(zn,(i,j)))\mu^n_{ij}, \sigma^n_{ij} = g_\psi(\text{context}(z^n, (i, j)))

    For a target image xx and rate-distortion trade-off multiplier λ>0\lambda > 0, the continuous latents zz and network parameters (θ,ψ)(\theta, \psi) are optimized jointly by minimizing the objective:

    Lθ,ψ(z)=∥x−fθ(Up(z))∥22−λ∑n=1Nlog⁡2Pψ(zn)L_{\theta,\psi}(z) = \|x - f_\theta(\text{Up}(z))\|_2^2 - \lambda \sum_{n=1}^N \log_2 P_\psi(z^n)

    Following optimization, the network parameters θ\theta and ψ\psi are quantized, entropy-coded using factorized zero-mean Laplace distributions scaled by their empirical standard deviations, and stored in the bitstream alongside the range-coded quantized latents.

  2. Knowl 2 — Quantization-Aware Two-Stage Optimization in C3

    model/method

    C3 trains latents zz and network weights (θ,ψ)(\theta, \psi) using a two-stage quantization-aware optimization strategy designed to improve convergence while maintaining fidelity to discrete quantization:

    • Stage 1 (Soft-Rounding with Kumaraswamy Noise): Latents are transformed using an invertible soft-rounding function sT(⋅)s_T(\cdot) before and after noise addition:

    z~=sT(sT(z)+ukum)\tilde{z} = s_T(s_T(z) + u_{\text{kum}})

    where T>0T > 0 is a temperature parameter that is annealed downward over optimization such that:

    lim⁡T→0sT(sT(z)+u)=⌊⌊z⌉+u⌉=⌊z⌉\lim_{T \to 0} s_T(s_T(z) + u) = \lfloor \lfloor z \rceil + u \rceil = \lfloor z \rceil

    Instead of standard uniform noise, ukumu_{\text{kum}} is sampled from a Kumaraswamy distribution on [0,1][0, 1] whose shape parameters transition from a low-variance peaked distribution at the beginning of Stage 1 to a uniform distribution at the end. Stage 1 employs an Adam optimizer with a cosine learning rate decay schedule. Quantization step sizes smaller than 11 are used (with rescaled soft-rounding) to prevent large magnitude inputs to downstream networks.

    • Stage 2 (Quantized Forward Pass with Soft-Rounding Gradients): Latents are hard-quantized ⌊z⌉\lfloor z \rceil during the forward pass. Backward gradients ∇zL\nabla_z L are computed using the derivative of the soft-rounding function evaluated at a low temperature rather than straight-through estimation (STE). Stage 2 starts at a reduced learning rate and adaptively decreases the learning rate whenever the rate-distortion loss fails to decrease over a specified iteration window.
  3. Knowl 3 — C3 Extensions for Video Compression

    model/method

    To compress video sequences without amortized encoders or explicit optical flow networks, C3 introduces four spatio-temporal adaptations:

    1. 3D Latent Grids and 3D Context: Latents z=(z1,…,zN)z = (z^1, \dots, z^N) are extended to 3D grids of shape (t,h,w),(t/2,h/2,w/2),…(t, h, w), (t/2, h/2, w/2), \dots, where tt is the number of frames and h,wh, w are frame dimensions. The context for entropy network gψg_\psi is structured as a 3D causal neighborhood around zτijnz^n_{\tau i j} in temporal frame τ\tau.
    2. Video Patch Partitioning: High-definition video sequences are split into spatio-temporal patches ranging from (30,180,240)(30, 180, 240) to (75,270,320)(75, 270, 320) (frames ×\times height ×\times width), fitting an independent C3 instance per patch to stay within GPU memory limits during gradient optimization.
    3. Expanded Spatial Context: The spatial context window in previous frame τ−1\tau - 1 is widened up to 65 latent pixels to capture keypoint displacement in fast-motion scenes.
    4. Custom Learned Masking: To avoid a linear parameter explosion in the entropy network from wide contexts, a custom mask is applied: a small 3D causal mask centered at the target latent in frame τ\tau, and a compact rectangular sub-window in frame τ−1\tau - 1 whose spatial offset is learned during encoding.
  4. Knowl 4 — C3 Architectural and Parametric Enhancements

    model/method

    C3 incorporates several architecture-level modifications over Cool-chic to increase model capacity within small parameter budgets:

    1. GELU Non-linearities: Rectified Linear Units (ReLU) across the synthesis and entropy networks are replaced by Gaussian Error Linear Units (GELU), improving reconstruction quality without increasing MAC count.
    2. Shifted Log-Scale Entropy Output: The raw network output predicting the Laplace scale parameter σ\sigma is shifted prior to exponentiation, stabilizing early optimization and enabling larger initialization scale values.
    3. Cross-Resolution Context and Resolution Conditioning: The entropy model optionally incorporates context from coarser latent grids P(zn∣zn−1)P(z^n \mid z^{n-1}) or uses Feature-wise Linear Modulation (FiLM) layers / separate networks per latent grid to provide resolution-dependent density modeling.
    4. Instance Adaptivity: For each image or video patch, hyperparameter sweeps (e.g., selecting whether to disable the highest-resolution latent grid at low bitrates) can be evaluated to pick the Pareto-optimal configuration on a per-instance basis.
  5. Knowl 5 — Image Rate-Distortion Performance and Decoding Complexity

    empirical result

    On the CLIC2020 professional validation set and Kodak image benchmarks, C3 achieves rate-distortion performance competitive with state-of-the-art classical and neural codecs while maintaining decoding arithmetic complexity below 3k MACs/pixel3\text{k MACs/pixel}:

    • On CLIC2020, fixed C3 matches VTM (the reference implementation of H.266/VVC) with a BD-rate of −0.1%-0.1\% and outperforms Cool-chic v2 by −20.8%-20.8\% BD-rate. Instance-adaptive C3 outperforms VTM with a BD-rate of −2.9%-2.9\%.
    • On Kodak, C3 achieves an RD trade-off superior to Cool-chic v2, BPG (HEVC), and neural baseline BMS. Neural codecs achieving comparable or superior BD-rates to C3 (such as EVC-S/M/L, CST, and MLIC+) require between 10410^4 and 106 MACs/pixel10^6\text{ MACs/pixel}, representing one to two orders of magnitude higher decoding compute than C3.
  6. Knowl 6 — Video Rate-Distortion Performance on UVG-1k Benchmark

    empirical result

    On the UVG-1k video benchmark (7 HD 1080×19201080 \times 1920 sequences, 3900 frames total evaluated in RGB PSNR), C3 achieves rate-distortion performance matching the Video Compression Transformer (VCT) while requiring 4.4k MACs/pixel4.4\text{k MACs/pixel} for decoding. This decoding complexity is less than 0.1%0.1\% of VCT's decoding compute (>106 MACs/pixel>10^6\text{ MACs/pixel}).

    Among neural video codecs based on single-instance overfitting, C3 outperforms FFNeRV and ranks second behind HiNeRV in RD performance. Whereas HiNeRV scales its model size with bitrate (requiring 87k87\text{k} to 1.2M MACs/pixel1.2\text{M MACs/pixel} across its RD curve), C3 maintains a constant decoding complexity of 4.4k MACs/pixel4.4\text{k MACs/pixel}.

  7. Knowl 7 — Kodak Ablation of C3 Optimization and Architectural Components

    data/table

    Ablation results on the Kodak image benchmark isolate the performance impact of each optimization and architectural modification introduced in C3.

    Sequential removal of improvements from C3 (adaptive) results in the following BD-rate degradations relative to C3 Adaptive:

    Model Variant BD-rate vs. C3 Adaptive (%)
    C3 (adaptive) 0.0%
    C3 (fixed architecture) 2.2%
    ×\times Quantization step <1< 1 2.6%
    ×\times Adaptive lr (stage 2) 3.4%
    ×\times Shift log-scale 4.2%
    ×\times GELU (revert to ReLU) 12.6%
    ×\times Kumaraswamy noise (revert to Uniform) 23.6%
    ×\times Soft-rounding (revert to noise / STE) 39.8%

    Individual feature knockouts from fixed C3 show that soft-rounding (+22.18% BD-rate), Kumaraswamy noise (+3.90% BD-rate), and GELU (+3.27% BD-rate) drive the largest individual rate-distortion improvements, followed by shifted log-scale (+0.87%), Stage 2 adaptive learning rate (+0.68%), and sub-unit quantization steps (+0.40%).

  8. Knowl 8 — Computational Runtimes, Decoding Latency, and Encoding Bottlenecks

    limitation

    C3 achieves low arithmetic decoding complexity and sub-second CPU decoding latency, but suffers from long encoding times due to per-instance gradient optimization:

    • Decoding Latency: On a single CPU core (Intel Xeon Platinum 2GHz), full decoding of a 768×512768 \times 512 image requires <100 ms< 100\text{ ms} (∼55 ms\sim 55\text{ ms} for autoregressive entropy model rollout, ∼30 ms\sim 30\text{ ms} for synthesis network upsampling and convolutions), excluding range decoding.
    • Image Encoding Runtime: On an NVIDIA V100 GPU, optimizing a single 1370×20481370 \times 2048 CLIC image requires 22s (smallest architecture) to 48s (largest architecture) per 1,000 iterations for up to 110,000 optimization iterations.
    • Video Encoding Runtime: Video patch optimization on a V100 GPU requires 29s per 1,000 iterations for small patches (30×180×24030 \times 180 \times 240) and 457s per 1,000 iterations for large patches (75×270×32075 \times 270 \times 320).
    • Sequential Entropy Decoding: Autoregressive spatial-temporal context modeling is intrinsically causal and sequential, limiting GPU hardware parallelization during decoding without custom implementations such as wavefront processing.

Coverage note — Appendix-specific implementation details, exhaustive baseline configuration setups, and extended qualitative gallery figures were omitted in favor of the core architectural, algorithmic, empirical, and ablation contributions.

References

  1. 1.Anne Aaron, Zhi Li, Megha Manohara, Jan De Cock, and David Ronca. Per-title encode optimization. https://netflixtechblog.com/per-title-encode-optimization-7e99442b62a2, 2015. [Online; accessed 26-Feb-2021]. 8
  2. 2.Eirikur Agustsson and Lucas Theis. Universally quantized neural compression. In Proceedings of the 34th International Conference on Neural Information Processing Systems, Red Hook, NY, USA, 2020. Curran Associates Inc. 3, 17
  3. 3.Eirikur Agustsson, David Minnen, Nick Johnston, Johannes Balle, Sung Jin Hwang, and George Toderici. Scale-space flow for end-to-end optimized video compression. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8503–8512, 2020. 5
  4. 4.Yunpeng Bai, Chao Dong, Cairong Wang, and Chun Yuan. Ps-nerv: Patch-wise stylized neural representations for videos. In 2023 IEEE International Conference on Image Processing (ICIP), pages 41–45. IEEE, 2023. 5
  5. 5.J Balle, V Laparra, and E P Simoncelli. End-to-end optimized image compression. In Int’l Conf on Learning Representations (ICLR), Toulon, France, 2017. Available at http://arxiv.org/abs/1611.01704. 1, 3
  6. 6.Johannes Balle, David Minnen, Saurabh Singh, Sung Jin Hwang, and Nick Johnston. Variational image compression with a scale hyperprior. In International Conference on Learning Representations, 2018. 1, 6, 15, 27
  7. 7.Jean Bégaint, Fabien Racapé, Simon Feltman, and Akshay Pushparaja. Compressai: a pytorch library and evaluation platform for end-to-end compression research. arXiv preprint arXiv:2011.03029, 2020. 28
  8. 8.Fabrice Bellard. Bpg image format. URL https://bellard.org/bpg, 1(2):1, 2015. 2, 6
  9. 9.Gisle Bjontegaard. Calculation of average psnr differences between rd-curves. ITU SG16 Doc. VCEG-M33, 2001. 2, 26
  10. 10.James Bradbury, Roy Frostig, Peter Hawkins, Matthew James Johnson, Chris Leary, Dougal Maclaurin, George Necula, Adam Paszke, Jake VanderPlas, Skye Wanderman-Milne, and Qiao Zhang. JAX: composable transformations of Python+NumPy programs, 2018. 27
  11. 11.Benjamin Bross, Ye-Kui Wang, Yan Ye, Shan Liu, Jianle Chen, Gary J Sullivan, and Jens-Rainer Ohm. Overview of the versatile video coding (vvc) standard and its applications. IEEE Transactions on Circuits and Systems for Video Technology, 31(10):3736–3764, 2021. 2, 6, 28
  12. 12.Joaquim Campos, Meierhans Simon, Abdelaziz Djelouah, and Christopher Schroers. Content adaptive optimization for neural image compression. In CVPR Workshop and Challenge on Learned Image Compression, 2019. 5
  13. 13.Lorenzo Catania and Dario Allegra. Nif: A fast implicit image compression with bottleneck layers and modulated sinusoidal activations. In Proceedings of the 31st ACM International Conference on Multimedia, pages 9022–9031, 2023. 5
  14. 14.Hao Chen, Bo He, Hanyu Wang, Yixuan Ren, Ser Nam Lim, and Abhinav Shrivastava. Nerv: Neural representations for videos. Advances in Neural Information Processing Systems, 34:21557–21568, 2021. 5, 28
  15. 15.Hao Chen, Matthew Gwilliam, Ser-Nam Lim, and Abhinav Shrivastava. Hnerv: A hybrid neural representation for videos. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10270–10279, 2023. 5, 28, 38
  16. 16.Zhengxue Cheng, Heming Sun, Masaru Takeuchi, and Jiro Katto. Learned image compression with discretized gaussian mixture likelihoods and attention modules. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 7939–7948, 2020. 1, 6, 27
  17. 17.Gordon Clare, Félix Henry, and Stéphane Pateux. Wavefront parallel processing for hevc encoding and decoding. document JCTVC-F274, 2011. 8
  18. 18.Bharath Bhushan Damodaran, Muhammet Balcilar, Franck Galpin, and Pierre Hellier. Rqat-inr: Improved implicit neural image compression. In 2023 Data Compression Conference (DCC), pages 208–217. IEEE, 2023. 5
  19. 19.Thomas Davies, Derek Nowrouzezahrai, and Alec Jacobson. On the effectiveness of weight-encoded neural implicit 3d shapes. arXiv preprint arXiv:2009.09808, 2020. 5
  20. 20.Emilien Dupont, Adam Goliński, Milad Alizadeh, Yee Whye Teh, and Arnaud Doucet. Coin: Compression with implicit neural representations. arXiv preprint arXiv:2103.03123, 2021. 1, 5, 28
  21. 21.Emilien Dupont, Hrushikesh Loya, Milad Alizadeh, Adam Golinski, Yee Whye Teh, and Arnaud Doucet. Coin++: Neural compression across modalities. Transactions on Machine Learning Research, 2022. 5, 8, 28
  22. 22.The fvcore contributors. fvcore. https://github.com/facebookresearch/fvcore. 27
  23. 23.Harry Gao, Weijie Gan, Zhixin Sun, and Ulugbek S Kamilov. Sinco: A novel structural regularizer for image compression using implicit neural representations. In ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 1–5. IEEE, 2023. 5
  24. 24.Sharath Girish, Abhinav Shrivastava, and Kamal Gupta. Shacira: Scalable hash-grid compression for implicit neural representations. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 17513–17524, 2023. 5
  25. 25.Carlos Gomes, Roberto Azevedo, and Christopher Schroers. Video compression with entropy-constrained neural representations. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 18497–18506, 2023. 5
  26. 26.Cameron Gordon, Shin-Fang Chng, Lachlan MacDonald, and Simon Lucey. On quantizing implicit neural representations. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 341–350, 2023. 5
  27. 27.Tiansheng Guo, Jing Wang, Ze Cui, Yihui Feng, Yunying Ge, and Bo Bai. Variable rate image compression with content adaptive optimization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, pages 122–123, 2020. 5
  28. 28.Zongyu Guo, Gergely Flamich, Jiajun He, Zhibo Chen, and José Miguel Hernández-Lobato. Compression with bayesian implicit neural representations. arXiv preprint arXiv:2305.19185, 2023. 5
  29. 29.Wang Guo-Hua, Jiahao Li, Bin Li, and Yan Lu. EVC: Towards real-time neural image compression with mask decay. In The Eleventh International Conference on Learning Representations, 2023. 2, 6, 27
  30. 30.Dailan He, Yaoyan Zheng, Baocheng Sun, Yan Wang, and Hongwei Qin. Checkerboard context model for efficient learned image compression. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 14771–14780, 2021. 6
  31. 31.Dailan He, Ziming Yang, Weikun Peng, Rui Ma, Hongwei Qin, and Yan Wang. Elic: Efficient learned image compression with unevenly grouped space-channel contextual adaptive coding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5718–5727, 2022. 1, 6, 28
  32. 32.Jiajun He, Gergely Flamich, Zongyu Guo, and José Miguel Hernández-Lobato. Recombiner: Robust and enhanced compression with bayesian implicit neural representations. arXiv preprint arXiv:2309.17182, 2023. 5, 28
  33. 33.Dan Hendrycks and Kevin Gimpel. Gaussian error linear units (gelus). arXiv preprint arXiv:1606.08415, 2016. 4, 13
  34. 34.Langwen Huang and Torsten Hoefler. Compressing multidimensional weather and climate data into neural networks. In The Eleventh International Conference on Learning Representations, 2023. 5
  35. 35.Berivan Isik, Philip A Chou, Sung Jin Hwang, Nick Johnston, and George Toderici. Lvac: Learned volumetric attribute compression for point clouds using coordinate based networks. Frontiers in Signal Processing, 2:1008812, 2022. 5
  36. 36.Itseez. Open source computer vision library. https://github.com/opencv/opencv, 2015. 19
  37. 37.ITU-T. Recommendation ITU-T T.81: Information technology – Digital compression and coding of continuous-tone still images – Requirements and guidelines, 1992. 4
  38. 38.Wei Jiang, Jiayu Yang, Yongqi Zhai, Peirong Ning, Feng Gao, and Ronggang Wang. Mlic: Multi-reference entropy model for learned image compression. In Proceedings of the 31st ACM International Conference on Multimedia, page 7618–7627, New York, NY, USA, 2023. Association for Computing Machinery. 1, 6, 27
  39. 39.Nick Johnston, Elad Eban, Ariel Gordon, and Johannes Ballé. Computationally efficient neural image compression. arXiv preprint arXiv:1912.08771, 2019. 6
  40. 40.Kodak. Kodak Dataset. http://r0k.us/graphics/kodak/, 1991. 6
  41. 41.P. Kumaraswamy. A generalized probability density function for double-bounded random processes. Journal of Hydrology, 46(1):79–88, 1980. 4, 18
  42. 42.Ho Man Kwan, Ge Gao, Fan Zhang, Andrew Gower, and David Bull. Hinerv: Video compression with hierarchical encoding based neural representation, 2023. 5, 7, 8, 28, 38
  43. 43.Théo Ladune, Pierrick Philippe, Félix Henry, Gordon Clare, and Thomas Leguay. Cool-chic: Coordinate-based low complexity hierarchical image codec. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 13515–13522, 2023. 1, 2, 6, 13, 15, 16, 17, 27
  44. 44.Luca A Lanzendörfer and Roger Wattenhofer. Siamese siren: Audio compression with implicit neural representations. arXiv preprint arXiv:2306.12957, 2023. 5
  45. 45.Hoang Le, Liang Zhang, Amir Said, Guillaume Sautiere, Yang Yang, Pranav Shrestha, Fei Yin, Reza Pourreza, and Auke Wiggers. Mobilecodec: neural inter-frame video compression on mobile devices. In Proceedings of the 13th ACM Multimedia Systems Conference, pages 324–330, 2022. 1, 6
  46. 46.Jaeho Lee, Jihoon Tack, Namhoon Lee, and Jinwoo Shin. Meta-learning sparse implicit neural representations. Advances in Neural Information Processing Systems, 34:11769–11780, 2021. 5
  47. 47.Joo Chan Lee, Daniel Rho, Jong Hwan Ko, and Eunbyung Park. Ffnerv: Flow-guided frame-wise neural representations for videos. 2023. 5, 7, 8
  48. 48.Thomas Leguay, Théo Ladune, Pierrick Philippe, Gordon Clare, and Félix Henry. Low-complexity overfitted neural image codec. arXiv preprint arXiv:2307.12706, 2023. 2, 6, 13, 16, 27, 33
  49. 49.Jiahao Li, Bin Li, and Yan Lu. Deep contextual video compression. Advances in Neural Information Processing Systems, 34:18114–18125, 2021. 7, 28
  50. 50.Jiahao Li, Bin Li, and Yan Lu. Hybrid spatial-temporal entropy modelling for neural video compression. In Proceedings of the 30th ACM International Conference on Multimedia, pages 1503–1511, 2022. 7
  51. 51.Jiahao Li, Bin Li, and Yan Lu. Neural video compression with diverse contexts. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 22616–22626, 2023. 7
  52. 52.Lingzhi Li, Zhen Shen, Zhongshu Wang, Li Shen, and Liefeng Bo. Compressing volumetric radiance fields to 1 mb. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4222–4231, 2023. 5
  53. 53.Zizhang Li, Mengmeng Wang, Huaijin Pi, Kechun Xu, Jianbiao Mei, and Yong Liu. E-nerv: Expedite neural video representation with disentangled spatial-temporal context. In European Conference on Computer Vision, pages 267–284. Springer, 2022. 5
  54. 54.Guo Lu, Chunlei Cai, Xiaoyun Zhang, Li Chen, Wanli Ouyang, Dong Xu, and Zhiyong Gao. Content adaptive and error propagation aware deep video compression. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part II 16, pages 456–472. Springer, 2020. 5
  55. 55.Yuzhe Lu, Kairong Jiang, Joshua A Levine, and Matthew Berger. Compressive neural representations of volumetric scalar fields. In Computer Graphics Forum, pages 135–146. Wiley Online Library, 2021. 5
  56. 56.Bruce D Lucas and Takeo Kanade. An iterative image registration technique with an application to stereo vision. In IJCAI’81: 7th international joint conference on Artificial intelligence, pages 674–679, 1981. 19
  57. 57.Yue Lv, Jinxi Xiang, Jun Zhang, Wenming Yang, Xiao Han, and Wei Yang. Dynamic low-rank instance adaptation for universal neural image compression. In Proceedings of the 31st ACM International Conference on Multimedia, pages 632–642, 2023. 5
  58. 58.Shishira R Maiya, Sharath Girish, Max Ehrlich, Hanyu Wang, Kwot Sin Lee, Patrick Poirson, Pengxiang Wu, Chen Wang, and Abhinav Shrivastava. Nirvana: Neural implicit representations of videos with adaptive networks and autoregressive patch-wise modeling. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 14378–14387, 2023. 5
  59. 59.Matteo Mancini, Derek K Jones, and Marco Palombo. Lossy compression of multidimensional medical images using sinusoidal activation networks: An evaluation study. In International Workshop on Computational Diffusion MRI, pages 26–37. Springer, 2022. 5
  60. 60.Fabian Mentzer, George Toderici, David Minnen, Sergi Caelles, Sung Jin Hwang, Mario Lucic, and Eirikur Agustsson. VCT: A video compression transformer. In Advances in Neural Information Processing Systems, 2022. 1, 2, 6, 7, 8, 26, 28
  61. 61.Alexandre Mercat, Marko Viitanen, and Jarno Vanne. Uvg dataset: 50/120fps 4k sequences for video codec analysis and development. In Proceedings of the 11th ACM Multimedia Systems Conference, pages 297–302, 2020. 2, 7
  62. 62.Yu Mikami, Chihiro Tsutake, Keita Takahashi, and Toshiaki Fujii. An efficient image compression method based on neural network: An overfitting approach. In 2021 IEEE International Conference on Image Processing (ICIP), pages 2084–2088. IEEE, 2021. 5
  63. 63.David Minnen and Nick Johnston. Advancing the rate-distortion-computation frontier for neural image compression. In 2023 IEEE International Conference on Image Processing (ICIP), pages 2940–2944. IEEE, 2023. 7
  64. 64.David Minnen, Johannes Ballé, and George D Toderici. Joint autoregressive and hierarchical priors for learned image compression. Advances in neural information processing systems, 31, 2018. 1, 27, 28
  65. 65.G. Nigel and N. Martin. Range encoding: An algorithm for removing redundancy from a digitized message. In Video & Data Recording Conference, 1979. 3
  66. 66.Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. Pytorch: An imperative style, high-performance deep learning library. In Advances in Neural Information Processing Systems 32, pages 8024–8035. Curran Associates, Inc., 2019. 27
  67. 67.Ethan Perez, Florian Strub, Harm de Vries, Vincent Dumoulin, and Aaron C. Courville. Film: Visual reasoning with a general conditioning layer. In AAAI, 2018. 4
  68. 68.Tuan Pham, Yibo Yang, and Stephan Mandt. Autoencoding implicit neural representations for image compression. In ICML 2023 Workshop Neural Compression: From Information Theory to Applications, 2023. 5
  69. 69.Juan Ramirez and Jose Gallego-Posada. L0onie: Compressing coins with l0-constraints, 2022. 5
  70. 70.Oren Rippel, Alexander G Anderson, Kedar Tatwawadi, Sanjay Nair, Craig Lytle, and Lubomir Bourdev. Elf-vc: Efficient learned flexible-rate video coding. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 14479–14488, 2021. 6, 28
  71. 71.Jonathan Schwarz and Yee Whye Teh. Meta-learning sparse compression networks. Transactions on Machine Learning Research, 2022. 5, 8, 28
  72. 72.Jonathan Richard Schwarz, Jihoon Tack, Yee Whye Teh, Jaeho Lee, and Jinwoo Shin. Modality-agnostic variational compression of implicit neural representations. In Proceedings of the 40th International Conference on Machine Learning. JMLR.org, 2023. 5, 28
  73. 73.Armin Sheibanifard and Hongchuan Yu. A novel implicit neural representation for volume data. Applied Sciences, 13(5):3242, 2023. 5
  74. 74.Xihua Sheng, Jiahao Li, Bin Li, Li Li, Dong Liu, and Yan Lu. Temporal context mining for learned video compression. IEEE Transactions on Multimedia, 2022. 7
  75. 75.Yibo Shi, Yunying Ge, Jing Wang, and Jue Mao. Alphavc: High-performance and efficient learned video compression. In European Conference on Computer Vision, pages 616–631. Springer, 2022. 6
  76. 76.Yannick Strümpler, Janis Postels, Ren Yang, Luc Van Gool, and Federico Tombari. Implicit neural representations for image compression. In European Conference on Computer Vision, pages 74–91. Springer, 2022. 5, 8, 28
  77. 77.Gary J Sullivan, Jens-Rainer Ohm, Woo-Jin Han, and Thomas Wiegand. Overview of the high efficiency video coding (hevc) standard. IEEE Transactions on circuits and systems for video technology, 22(12):1649–1668, 2012. 2, 4, 7, 8, 28
  78. 78.Towaki Takikawa, Alex Evans, Jonathan Tremblay, Thomas Müller, Morgan McGuire, Alec Jacobson, and Sanja Fidler. Variable bitrate neural fields. In ACM SIGGRAPH 2022 Conference Proceedings, pages 1–9, 2022. 5
  79. 79.Lucas Theis, Wenzhe Shi, Andrew Cunningham, and Ferenc Huszár. Lossy image compression with compressive autoencoders. arXiv preprint arXiv:1703.00395, 2017. 1
  80. 80.George Toderici, Wenzhe Shi, Radu Timofte, Lucas Theis, Johannes Balle, Eirikur Agustsson, Nick Johnston, and Fabian Mentzer. Workshop and challenge on learned image compression (clic2020). In CVPR, 2020. 2, 6
  81. 81.Ties Van Rozendaal, Johann Brehmer, Yunfan Zhang, Reza Pourreza, Auke Wiggers, and Taco S Cohen. Instance-adaptive video compression: Improving neural codecs by training on the test set. arXiv preprint arXiv:2111.10302, 2021. 5, 28
  82. 82.Ties van Rozendaal, Iris AM Huijben, and Taco Cohen. Overfitting for fun and profit: Instance-adaptive data compression. In International Conference on Learning Representations, 2021. 5, 7
  83. 83.Ties van Rozendaal, Tushar Singhal, Hoang Le, Guillaume Sautiere, Amir Said, Krishna Buska, Anjuman Raha, Dimitris Kalatzis, Hitarth Mehta, Frank Mayer, et al. Mobilenvc: Real-time 1080p neural video compression on a mobile device. arXiv preprint arXiv:2310.01258, 2023. 1, 6
  84. 84.Adam Wieckowski, Jens Brandenburg, Tobias Hinz, Christian Bartnik, Valeri George, Gabriel Hege, Christian Helmrich, Anastasia Henkel, Christian Lehmann, Christian Stoffers, Ivan Zupancic, Benjamin Bross, and Detlev Marpe. Vvenc: An open and optimized vvc encoder implementation. In Proc. IEEE International Conference on Multimedia Expo Workshops (ICMEW), pages 1–2. 28
  85. 85.Jinxi Xiang, Kuan Tian, and Jun Zhang. Mimt: Masked image modeling transformer for video compression. In The Eleventh International Conference on Learning Representations, 2023. 7
  86. 86.Yiheng Xie, Towaki Takikawa, Shunsuke Saito, Or Litany, Shiqin Yan, Numair Khan, Federico Tombari, James Tompkin, Vincent Sitzmann, and Srinath Sridhar. Neural fields in visual computing and beyond. In Computer Graphics Forum, pages 641–676. Wiley Online Library, 2022. 1
  87. 87.Yibo Yang and Stephan Mandt. Computationally-efficient neural image compression with shallow decoders. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 530–540, 2023. 6
  88. 88.Yibo Yang, Robert Bamler, and Stephan Mandt. Improving inference for neural image compression. Advances in Neural Information Processing Systems, 33:573–584, 2020. 5
  89. 89.Yibo Yang, Stephan Mandt, Lucas Theis, et al. An introduction to neural data compression. Foundations and Trends® in Computer Graphics and Vision, 15(2):113–200, 2023. 1, 2, 6
  90. 90.Renjie Zou, Chunfeng Song, and Zhaoxiang Zhang. The devil is in the details: Window-based attention for image compression. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 17492–17501, 2022. 28

Citation

MLA
Kim, H., et al. “C3: High-performance and Low-complexity Neural Compression from a Single Image or Video”. arXiv, 2023, http://arxiv.org/abs/2312.02753v1.
APA
Kim, H., Bauer, M., Theis, L., Schwarz, J. R., & Dupont, E. (2023). C3: High-performance and low-complexity neural compression from a single image or video. arXiv. http://arxiv.org/abs/2312.02753v1
Chicago
Kim, H., M. Bauer, L. Theis, J. R. Schwarz, and E. Dupont. 2023. “C3: High-performance and Low-complexity Neural Compression from a Single Image or Video”. arXiv. http://arxiv.org/abs/2312.02753v1.
Harvard
Kim, H. et al. (2023) “C3: High-performance and low-complexity neural compression from a single image or video”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2312.02753v1.
Vancouver
1. Kim H, Bauer M, Theis L, Schwarz JR, Dupont E (2023) C3: High-performance and low-complexity neural compression from a single image or video. arXiv

BibTeX

@article{kim2023high,
  title = {C3: High-performance and low-complexity neural compression from a single image or video},
  author = {Kim, Hyunjik and Bauer, Matthias and Theis, Lucas and Schwarz, Jonathan Richard and Dupont, Emilien},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2312.02753v1},
  eprint = {2312.02753}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE