ELIC: Efficient Learned Image Compression with Unevenly Grouped Space-Channel Contextual Adaptive Coding
Dailan HeZiming YangWeikun PengRui MaHongwei QinYan Wang
Presents ELIC, a learned image compression architecture that combines uneven space-channel contextual coding with efficient transform design to achieve state-of-the-art rate-distortion performance alongside fast inference, preview decoding, and progressive decoding.
Modern digital workflows demand efficient image compression to reduce storage costs and network bandwidth without sacrificing visual fidelity. While artificial intelligence approaches have recently surpassed traditional compression formats in quality, their practical deployment has been severely hindered by excessive computational complexity and slow decoding speeds. Many advanced neural compression methods process data serially, creating significant latency bottlenecks that make high-volume or real-time application impractical. The article addresses this critical trade-off between compression performance and operational speed.
The main objective of the article is to design, evaluate, and demonstrate an efficient learned image compression framework named ELIC that achieves state-of-the-art compression quality while maintaining fast running speeds. The researchers evaluate whether reorganizing how image information is processed and simplifying the underlying neural network architecture can outperform leading industrial standards in both compression ratio and latency.
To accomplish this, the authors develop a space-channel context model that splits latent feature channels into unevenly sized groups, processing earlier high-information channels at fine granularity and later channels in larger, parallel chunks. They combine this with a parallel spatial context model to capture dependencies across multiple dimensions without slowing processing. Additionally, they replace traditional divisive normalization layers with stacked residual bottleneck blocks and build a lightweight thumbnail synthesizer for low-cost image previews. The approach was systematically evaluated on standard benchmark datasets, including Kodak and CLIC Professional, after training on 8,000 high-resolution images from ImageNet, comparing both compression efficiency and hardware latency against modern traditional codecs and leading learned alternatives.
The findings demonstrate substantial improvements across compression and latency metrics. First, ELIC outperforms the leading modern standard, Versatile Video Coding (VVC) in YUV 4:4:4 format, achieving a 7.88% bitrate reduction on the Kodak dataset while maintaining equal objective quality. When optimized specifically for visual similarity, it saves approximately 50% of the bitrate compared to VVC. Second, the uneven grouping strategy cuts the latency of adaptive entropy estimation roughly in half compared to conventional ten-slice even grouping. Third, the full ELIC model achieves total encoding and decoding latencies of approximately 42.4 milliseconds and 49.2 milliseconds per image on standard hardware, whereas prior top-performing serial models exceed 1,000 milliseconds. A streamlined version, ELIC-sm, further lowers decoding time to 27.8 milliseconds while maintaining performance superior to VVC. Finally, the dedicated thumbnail synthesizer generates preview images in about 3 microseconds—over 12 times faster than running the full reconstruction synthesizer.
These results indicate that deep learning-based compression has reached operational viability for commercial systems. By achieving symmetric encoding and decoding speeds under 50 milliseconds, ELIC removes the multi-second decoding bottleneck that previously prevented practical adoption. This enables high-throughput data pipelines, reduced cloud storage expenses, lower bandwidth consumption, and responsive browsing via rapid preview generation, all while matching or exceeding the best handcrafted compression standards.
Organizations evaluating next-generation image infrastructure should consider testing learned compression frameworks like ELIC for high-performance transmission and storage workflows. Engineering teams can choose between the standard ELIC model for maximum bitrate savings or ELIC-sm for lower latency constraints. For applications involving media galleries or progressive web loading, deploying the dedicated thumbnail synthesizer is recommended to eliminate unnecessary full-resolution compute overhead.
The findings carry high confidence within the evaluated testbeds, supported by reproducible quantitative benchmarks on standard image collections. However, decision-makers should note that the primary comparisons against VVC used the YUV 4:4:4 color format rather than the standard YUV 4:2:0 format commonly utilized in broadcast and video pipelines. Further evaluation across broader subsampled color formats and mobile edge hardware is advised before executing large-scale production migrations.
- Paper: Joint Autoregressive and Hierarchical Priors for Learned Image Compression, David Minnen et al. (2018). This seminal work establishes the joint autoregressive and hierarchical prior framework that ELIC directly builds upon and accelerates through uneven channel grouping.
- Paper: Variational image compression with a scale hyperprior, Johannes Ballé et al. (2018). This paper introduces the foundational scale hyperprior architecture for learned image compression that underpins modern entropy modeling.
- Paper: Learned Image Compression With Discretized Gaussian Mixture Likelihoods and Attention Modules, Zhengxue Cheng et al. (2020). It provides essential background on incorporating attention mechanisms and advanced likelihood formulations into deep end-to-end image compression models.
- Paper: End-to-end Optimized Image Compression, Johannes Ballé et al. (2016). It provides the foundational end-to-end rate-distortion optimization framework and generalized divisive normalization transforms used throughout neural image compression.
- Paper: ShuffleNet V2: Practical Guidelines for Efficient CNN Architecture Design, Ningning Ma et al. (2018). It presents foundational design principles for balancing channel splitting and computation speed, directly motivating efficient transform designs in fast neural networks.
No sufficiently relevant recommendations were found.
