Large Images are Gaussians: High-Quality Large Image Representation with Levels of 2D Gaussian Splatting
Lingting ZhuGuying LinJinnan ChenXinjie ZhangZhenchao JinZhao WangLequan Yu
Introduces a multi-level 2D Gaussian splatting framework that scales to large images by decoupling coarse and fine details, providing faster decoding and superior fidelity compared to implicit neural representations.
Representing high-resolution visual data efficiently is a critical challenge in fields such as satellite communication and telemedicine. While traditional continuous representations using implicit neural networks capture intricate details, they require prohibitive amounts of computing memory and suffer from slow decoding speeds when scaled to large images. Recent adaptations of 2D Gaussian Splatting offer faster rendering, but existing frameworks fail to maintain high image quality when handling the tens of millions of Gaussian points necessary for very large images.
The article demonstrates and evaluates a novel framework called Large Images are Gaussians. The primary objective is to enable high-fidelity, computationally efficient fitting of ultra-high-resolution images using a scaled 2D Gaussian Splatting architecture.
To achieve this, the approach introduces two fundamental innovations. First, it reconfigures the mathematical representation by directly optimizing the 2D covariance matrix with custom hardware acceleration, filtering out invalid values post-processing rather than using complex parameter decompositions. Second, it implements a two-stage hierarchical fitting mechanism where a small fraction of Gaussian points captures coarse low-frequency image foundations from a downsampled version, while the remaining majority learns the high-frequency residual details. The authors evaluated this framework on high-resolution benchmarks spanning 2K general visual images, 4K satellite scenes, and 9K medical histopathology slides, benchmarking performance against established implicit neural networks and existing Gaussian fitting methods.
The findings confirm substantial performance improvements across all evaluation datasets. Most importantly, the framework consistently improves image fidelity as the number of Gaussian points grows, whereas baseline Gaussian methods deteriorate under large point counts. On 4K satellite imagery, the proposed method achieves superior quality scores reaching 56.05 dB, compared to around 27.50 dB for the baseline Gaussian approach. On 9K medical images, it achieves quality scores up to 42.19 dB while lowering peak training memory consumption to approximately 20.26 GB, well below the prohibitive memory demands of neural network baselines. Additionally, the approach delivers high rendering speeds ranging from tens to hundreds of frames per second, far exceeding neural network alternatives.
These results demonstrate that 2D Gaussian Splatting is a viable, high-performance alternative to neural representations for very large images. By reducing training memory and retaining real-time rendering capabilities, the method significantly lowers operational computing costs and processing bottlenecks. However, when compared against specialized state-of-the-art radial basis neural fields under equivalent parameter counts, the method shows slightly lower reconstruction accuracy, highlighting that 2D Gaussian image representation remains in its early development.
Organizations handling large-scale imagery should consider adopting two-stage Gaussian representations for workflows demanding real-time rendering and constrained training memory. For future development, researchers and practitioners should prioritize integrating parameter compression techniques to reduce the total number of Gaussians required, bringing overall compression efficiency in line with leading neural representations.
- Paper: 2D Gaussian Splatting for Geometrically Accurate Radiance Fields, Binbin Huang et al. (2024). This introduces 2D Gaussian Splatting, the image-fitting representation LIG directly adapts, so it clarifies the primitives and optimization choices behind the source.
- Paper: Deep Generative Image Models using a Laplacian Pyramid of Adversarial Networks, Emily Denton et al. (2015). Its Laplacian-pyramid coarse-to-fine strategy provides useful grounding for LIG’s level-based reconstruction of low-frequency structure and fine image detail.
No sufficiently relevant recommendations were found.
