Large Images are Gaussians: High-Quality Large Image Representation with Levels of 2D Gaussian Splatting

Lingting ZhuGuying LinJinnan ChenXinjie ZhangZhenchao JinZhao WangLequan Yu

article2025AAAI39 citations

Introduces a multi-level 2D Gaussian splatting framework that scales to large images by decoupling coarse and fine details, providing faster decoding and superior fidelity compared to implicit neural representations.

Listen

Representing high-resolution visual data efficiently is a critical challenge in fields such as satellite communication and telemedicine. While traditional continuous representations using implicit neural networks capture intricate details, they require prohibitive amounts of computing memory and suffer from slow decoding speeds when scaled to large images. Recent adaptations of 2D Gaussian Splatting offer faster rendering, but existing frameworks fail to maintain high image quality when handling the tens of millions of Gaussian points necessary for very large images.

The article demonstrates and evaluates a novel framework called Large Images are Gaussians. The primary objective is to enable high-fidelity, computationally efficient fitting of ultra-high-resolution images using a scaled 2D Gaussian Splatting architecture.

To achieve this, the approach introduces two fundamental innovations. First, it reconfigures the mathematical representation by directly optimizing the 2D covariance matrix with custom hardware acceleration, filtering out invalid values post-processing rather than using complex parameter decompositions. Second, it implements a two-stage hierarchical fitting mechanism where a small fraction of Gaussian points captures coarse low-frequency image foundations from a downsampled version, while the remaining majority learns the high-frequency residual details. The authors evaluated this framework on high-resolution benchmarks spanning 2K general visual images, 4K satellite scenes, and 9K medical histopathology slides, benchmarking performance against established implicit neural networks and existing Gaussian fitting methods.

The findings confirm substantial performance improvements across all evaluation datasets. Most importantly, the framework consistently improves image fidelity as the number of Gaussian points grows, whereas baseline Gaussian methods deteriorate under large point counts. On 4K satellite imagery, the proposed method achieves superior quality scores reaching 56.05 dB, compared to around 27.50 dB for the baseline Gaussian approach. On 9K medical images, it achieves quality scores up to 42.19 dB while lowering peak training memory consumption to approximately 20.26 GB, well below the prohibitive memory demands of neural network baselines. Additionally, the approach delivers high rendering speeds ranging from tens to hundreds of frames per second, far exceeding neural network alternatives.

These results demonstrate that 2D Gaussian Splatting is a viable, high-performance alternative to neural representations for very large images. By reducing training memory and retaining real-time rendering capabilities, the method significantly lowers operational computing costs and processing bottlenecks. However, when compared against specialized state-of-the-art radial basis neural fields under equivalent parameter counts, the method shows slightly lower reconstruction accuracy, highlighting that 2D Gaussian image representation remains in its early development.

Organizations handling large-scale imagery should consider adopting two-stage Gaussian representations for workflows demanding real-time rendering and constrained training memory. For future development, researchers and practitioners should prioritize integrating parameter compression techniques to reduce the total number of Gaussians required, bringing overall compression efficiency in line with leading neural representations.

No sufficiently relevant recommendations were found.

Cover for Large Images are Gaussians: High-Quality Large Image Representation with Levels of 2D Gaussian Splatting

Abstract

While Implicit Neural Representations (INRs) have demonstrated significant success in image representation, they are often hindered by large training memory and slow decoding speed. Recently, Gaussian Splatting (GS) has emerged as a promising solution in 3D reconstruction due to its high-quality novel view synthesis and rapid rendering capabilities, positioning it as a valuable tool for a broad spectrum of applications. In particular, a GS-based representation, 2DGS, has shown potential for image fitting. In our work, we present \textbf{L}arge \textbf{I}mages are \textbf{G}aussians (\textbf{LIG}), which delves deeper into the application of 2DGS for image representations, addressing the challenge of fitting large images with 2DGS in the situation of numerous Gaussian points, through two distinct modifications: 1) we adopt a variant of representation and optimization strategy, facilitating the fitting of a large number of Gaussian points; 2) we propose a Level-of-Gaussian approach for reconstructing both coarse low-frequency initialization and fine high-frequency details. Consequently, we successfully represent large images as Gaussian points and achieve high-quality large image representation, demonstrating its efficacy across various types of large images. Code is available at {\href{this https URL}{this https URL}}.

Table of Contents

  • Introduction
  • Related Works
  • Methodology
  • Preliminaries
  • 2D Gaussians Formulation
  • Levels of 2D Gaussians
  • Experiments
  • Experimental Setup
  • Main Results
  • Ablation Studies
  • Further Analysis
  • Conclusion
  • Acknowledgments
  • References

Knowls

  1. Knowl 1 — Two-stage Level-of-Gaussian fitting separates coarse structure from image detail

    model/method

    Large Images are Gaussians (LIG) represents an image with two sets of 2D Gaussian splats. For target image II, level L0L_0 fits the image downsampled by a factor of two, providing a low-frequency initialization; level L1L_1 fits the residual between the full-resolution target and the upsampled rendering of L0L_0. The residual is min–max normalized to form the L1L_1 target, and the normalization extrema are saved for inference. The two levels are trained in sequence with mean-squared-error supervision and the same number of optimization steps per level; L0L_0 is frozen while L1L_1 is trained. At inference, the saved normalization values allow recovery of the residual scale so the coarse estimate and detail representation can be used to reconstruct the target.

    If N0N_0 and N1N_1 denote the Gaussian sets at the two levels, the allocation is ∣N∣=∣N0∣+∣N1∣=(1+r)∣N1∣|N|=|N_0|+|N_1|=(1+r)|N_1|, where ∣N∣|N| is the total number of Gaussians and r=∣N0∣/∣N1∣r=|N_0|/|N_1| is set to 0.1250.125 in the experiments. Thus, the coarser level uses fewer points. The targets are T0=Down⁡(I)T_0=\operatorname{Down}(I) and T1=Norm⁡(I−Up⁡(Render⁡(L0)))T_1=\operatorname{Norm}(I-\operatorname{Up}(\operatorname{Render}(L_0))), with downsampling and upsampling at the twofold resolution change. LIG uses two levels at all tested image resolutions. The authors report that freezing L0L_0 during L1L_1 training reduces the maximum number of simultaneously gradient-bearing points for a fixed total Gaussian count, reducing training memory; the second splatting pass adds some training and inference cost.

  2. Knowl 2 — LIG directly optimizes symmetric 2D covariance matrices and filters invalid splat contributions

    model/method

    Each image-space Gaussian has a 2D position μn\mu_n, a symmetric covariance matrix Σn∈R2×2\Sigma_n\in\mathbb{R}^{2\times2}, and a three-channel weighted color cn′c'_n. This gives eight parameters per Gaussian: two position values, three independent covariance values, and three color values. Unlike parameterizations that optimize a rotation and scale decomposition, LIG directly optimizes the three independent entries of Σn\Sigma_n using reimplemented CUDA forward and backward kernels.

    For pixel ii with center xi∈R2x_i\in\mathbb{R}^2, define the displacement din=xi−μnd_{in}=x_i-\mu_n and, for invertible Σn\Sigma_n, the scalar σin=12dinTΣn−1din\sigma_{in}=\tfrac12 d_{in}^{\mathsf T}\Sigma_n^{-1}d_{in}. The rendered color is

    Ci=∑n∈N: σin>0cn′exp⁡(−σin),C_i=\sum_{n\in\mathcal{N}:\,\sigma_{in}>0}c'_n\exp(-\sigma_{in}),

    where N\mathcal{N} is the set of splats considered for pixel ii, and Ci∈R3C_i\in\mathbb{R}^3. The weighted color formulation incorporates opacity into the color representation. Direct optimization does not guarantee that every covariance remains positive semidefinite, so the rendering rule excludes contributions with nonpositive σin\sigma_{in}. The paper notes that this pixelwise test does not identify every non-positive-semidefinite matrix, because it tests displacements from a finite pixel set; a strict 2D matrix test can instead be used when positive-semidefinite covariances are required.

  3. Knowl 3 — LIG improves large-image fitting quality over GaussianImage in the reported benchmarks

    data/table

    The table reports PSNR in dB, training memory in GB, and rendering speed in FPS for the STimage histopathology, FGF2 satellite, and DIV-HR general-image datasets. Gaussian counts in each GaussianImage or LIG row are listed in dataset order: STimage, FGF2, DIV-HR. A dash means no result was reported; the paper notes that grid-based INR fitting can be infeasible on large images because of memory requirements. The central pattern is that LIG's PSNR rises as the tested Gaussian counts increase, while GaussianImage's PSNR falls or remains nearly flat. LIG also uses less training memory than GaussianImage at the matched point-count settings shown. FPS is not uniformly higher for LIG, since its two levels require two splatting passes.

    Method (Gaussian counts) STimage (9K) FGF2 (4K) DIV-HR (2K)
    PSNR Mem. FPS PSNR Mem. FPS PSNR Mem. FPS
    SIREN – – – – – – 28.61 28.05 38.55
    Gauss – – – – – – 25.39 38.56 25.12
    WIRE – – – – – – 24.42 52.40 8.92
    FINER 17.74 71.27 12.81 21.91 74.82 12.35 34.42 80.01 9.84
    GaussianImage (3.5e7, 1e7, 5e5) 29.86 20.47 19.86 27.50 5.39 78.19 40.09 1.03 745.01
    GaussianImage (4.5e7, 1.2e7, 7e5) 29.33 23.56 18.95 27.53 5.64 69.24 35.12 1.09 635.30
    GaussianImage (5.5e7, 1.4e7, 9e5) 29.28 25.89 16.51 27.48 6.04 67.20 29.45 1.17 525.73
    LIG (3.5e7, 1e7, 5e5) 37.47 16.67 20.19 51.81 4.21 74.38 44.89 1.01 541.84
    LIG (4.5e7, 1.2e7, 7e5) 39.82 17.75 17.89 53.90 4.26 63.62 49.07 1.02 491.37
    LIG (5.5e7, 1.4e7, 9e5) 42.19 20.26 15.72 56.05 4.39 58.09 52.22 1.05 441.76

    For example, at the largest listed setting, LIG reaches 42.19 dB on STimage, 56.05 dB on FGF2, and 52.22 dB on DIV-HR, compared with GaussianImage's 29.28, 27.48, and 29.45 dB, respectively. LIG's corresponding training-memory values are 20.26, 4.39, and 1.05 GB, compared with 25.89, 6.04, and 1.17 GB.

  4. Knowl 4 — Ablations show gains from direct covariance optimization and the two-level design

    data/table

    This ablation compares GaussianImage, the LIG covariance-optimization variant without Level-of-Gaussian (LOG), and full LIG with LOG. Each method is evaluated at multiple Gaussian counts, with the same number of optimization iterations across the compared models. The covariance variant alone improves PSNR over GaussianImage across the listed settings; adding LOG further improves PSNR in every listed setting. The gains are especially pronounced on the higher-resolution STimage and FGF2 datasets.

    Dataset Gaussian count GaussianImage LIG without LOG LIG with LOG
    STimage (9K) 2.5e7 31.86 34.01 35.03
    STimage (9K) 3.5e7 29.86 34.02 37.47
    STimage (9K) 4.5e7 29.33 33.45 39.82
    STimage (9K) 5.5e7 29.28 32.68 42.19
    STimage (9K) 6.5e7 29.26 33.27 44.49
    FGF2 (4K) 6e6 28.05 47.07 47.17
    FGF2 (4K) 8e6 27.44 48.54 49.74
    FGF2 (4K) 1e7 27.50 49.65 51.81
    FGF2 (4K) 1.2e7 27.53 50.50 53.90
    FGF2 (4K) 1.4e7 27.48 51.35 56.05
    DIV-HR (2K) 5e5 40.09 43.21 44.89
    DIV-HR (2K) 6e5 38.13 43.66 47.07
    DIV-HR (2K) 7e5 35.12 43.47 49.07
    DIV-HR (2K) 8e5 32.00 42.97 50.82
    DIV-HR (2K) 9e5 29.45 42.30 52.22

    All reported values are PSNR in dB. The trend also shows that GaussianImage's quality decreases as point count increases in the listed settings, whereas the full LIG configuration improves across the tested counts.

  5. Knowl 5 — The coarse-level initialization improves the fit obtained from the second-level Gaussians

    data/table

    On STimage (9K), the authors compare full two-level fitting with a single-level fit that uses the same number of second-level Gaussians, thereby testing whether the first level's low-frequency initialization helps the final fit. For each point assignment, adding the first level increases PSNR but reduces FPS, reflecting the additional splatting work. The number pairs are (∣N0∣,∣N1∣)(|N_0|,|N_1|), where N0N_0 and N1N_1 are the coarse- and detail-level Gaussian sets.

    ∣N0∣|N_0| ∣N1∣|N_1| PSNR (dB) FPS
    0 30625000 34.12 25.78
    4375000 30625000 37.47 20.19
    0 39375000 33.81 21.04
    5625000 39375000 39.86 17.89
    0 48125000 33.22 17.34
    6875000 48125000 42.19 15.72

    At the three matched second-level counts, the two-level model gains 3.35, 6.05, and 8.97 dB, respectively, compared with the corresponding single-level model. The paper interprets this as evidence that the coarse fit makes the final residual target easier to learn.

  6. Knowl 6 — Datasets and training protocol used to evaluate LIG

    experimental setup

    LIG is evaluated on three image domains and resolutions: 15 histopathology images at 9K resolution from STimage, four 4K satellite images from the Full-resolution Gaofen-2 (FGF2) dataset, and 100 2K images from DIV-HR. The comparison metrics are PSNR for reconstruction quality, training memory in GB, and rendering speed in FPS. For LIG, both levels use the Adam optimizer with learning rate 0.0180.018; the implementation details report 30,000 training steps for each level. The representation is implemented with CUDA kernels and built on gsplat. Gaussian-count settings for GS-based methods vary by dataset and are reported as ordered triples for STimage, FGF2, and DIV-HR. The paper states that several INR results are omitted on the large datasets because the required grid-feature batch sizes make fitting infeasible; FINER is reported using different network sizes across datasets.

  7. Knowl 7 — More optimization iterations improve PSNR at increased training-time cost

    data/table

    The iteration study measures the quality–time trade-off for LIG on STimage (9K) and FGF2 (4K). PSNR rises with iteration count for both datasets, while training time also increases. The authors identify 3×1053\times10^5 iterations as a compromise used in their experiments. The Gaussian point counts for the two datasets are reported as (4.5e7,1.4e7)(4.5\mathrm{e}7,1.4\mathrm{e}7).

    Dataset Metric 1e51\mathrm{e}5 2e52\mathrm{e}5 3e53\mathrm{e}5 4e54\mathrm{e}5 5e55\mathrm{e}5
    STimage (9K) PSNR (dB) 37.68 39.12 39.82 40.38 40.76
    STimage (9K) Training time (s) 1825.68 3446.74 4777.65 6523.39 8028.46
    FGF2 (4K) PSNR (dB) 53.22 55.11 56.05 56.70 57.20
    FGF2 (4K) Training time (s) 670.15 1306.56 1926.32 2562.91 3186.11
  8. Knowl 8 — NeuRBF obtains higher FGF2 PSNR at a similar parameter count and rendering speed

    empirical result

    In the reported FGF2 comparison, LIG uses 2 million Gaussian points, corresponding to 16,000,000 optimized parameters. NeuRBF uses 17,631,486 parameters, a similar total. LIG achieves 41.31 dB PSNR at 168.25 FPS; NeuRBF achieves 44.69 dB at 156.38 FPS. Thus, in this matched-parameter comparison, NeuRBF has higher reconstruction quality while retaining a similar rendering speed. These figures qualify LIG's broader gains over the GaussianImage baseline: they do not establish that LIG outperforms all INR methods.

  9. Knowl 9 — The paper identifies Gaussian count and compression as unresolved limitations

    limitation

    LIG demonstrates image fitting with large numbers of Gaussian points, but the paper does not show that this representation achieves image compression quality comparable to state-of-the-art implicit neural representations. The authors identify reducing the number of Gaussians while retaining high image quality as an important next step for large-image compression. They also characterize Gaussian-splatting-based image representation as an early-stage area; whether efficient INR techniques such as hash coding can be applied to 2D Gaussian splatting remains unresolved.

Coverage note — Qualitative patch illustrations are omitted because they do not add a distinct quantitative finding beyond the benchmark and ablation results captured here; no other substantial contributed material is deliberately omitted.

References

  1. 1.Barron, J. T.; Mildenhall, B.; Tancik, M.; Hedman, P.; Martin-Brualla, R.; and Srinivasan, P. P. 2021. Mip-nerf: A multiscale representation for anti-aliasing neural radiance fields. In Proceedings of the IEEE/CVF international conference on computer vision, 5855–5864.
  2. 2.Bengio, Y.; Courville, A.; and Vincent, P. 2013. Representation learning: A review and new perspectives. IEEE transactions on pattern analysis and machine intelligence, 35(8): 1798–1828.
  3. 3.Chen, H.; Gwilliam, M.; Lim, S.-N.; and Shrivastava, A. 2023a. Hnerv: A hybrid neural representation for videos. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 10270–10279.
  4. 4.Chen, H.; He, B.; Wang, H.; Ren, Y.; Lim, S. N.; and Shrivastava, A. 2021. Nerv: Neural representations for videos. Advances in Neural Information Processing Systems, 34: 21557–21568.
  5. 5.Chen, J.; Li, C.; Zhang, J.; Zhu, L.; Huang, B.; Chen, H.; and Lee, G. H. 2024a. Generalizable Human Gaussians from Single-View Image. arXiv preprint arXiv:2406.06050.
  6. 6.Chen, J.; Zhou, M.; Wu, W.; Zhang, J.; Li, Y.; and Li, D. 2024b. STimage-1K4M: A histopathology image-gene expression dataset for spatial transcriptomics. arXiv:2406.06393.
  7. 7.Chen, Y.; Liu, S.; and Wang, X. 2021. Learning continuous image representation with local implicit image function. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 8628–8638.
  8. 8.Chen, Z.; Li, Z.; Song, L.; Chen, L.; Yu, J.; Yuan, J.; and Xu, Y. 2023b. Neurbf: A neural fields representation with adaptive radial basis functions. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 4182–4194.
  9. 9.Chen, Z.; and Zhang, H. 2019. Learning implicit fields for generative shape modeling. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 5939–5948.
  10. 10.De Sanctis, M.; Cianca, E.; Araniti, G.; Bisio, I.; and Prasad, R. 2015. Satellite communications supporting internet of remote things. IEEE Internet of Things Journal, 3(1): 113–123.
  11. 11.Dupont, E.; Golinski, A.; Alizadeh, M.; Teh, Y. W.; and Doucet, A. 2021. Coin: Compression with implicit neural representations. arXiv preprint arXiv:2103.03123.
  12. 12.Hu, J.; Xia, B.; Chen, B.; Yang, W.; and Zhang, L. 2024. GaussianSR: High Fidelity 2D Gaussian Splatting for Arbitrary-Scale Image Super-Resolution. arXiv preprint arXiv:2407.18046.
  13. 13.Huang, B.; Yu, Z.; Chen, A.; Geiger, A.; and Gao, S. 2024. 2d gaussian splatting for geometrically accurate radiance fields. In ACM SIGGRAPH 2024 Conference Papers, 1–11.
  14. 14.Kerbl, B.; Kopanas, G.; Leimkuhler, T.; and Drettakis, G. 2023. 3D Gaussian Splatting for Real-Time Radiance Field Rendering. ACM Trans. Graph., 42(4): 139–1.
  15. 15.Kerbl, B.; Meuleman, A.; Kopanas, G.; Wimmer, M.; Lanvin, A.; and Drettakis, G. 2024. A hierarchical 3d gaussian representation for real-time rendering of very large datasets. ACM Transactions on Graphics (TOG), 43(4): 1–15.
  16. 16.Khayam, S. A. 2003. The discrete cosine transform (DCT): theory and application. Michigan State University, 114(1): 31.
  17. 17.Kingma, D. P.; and Ba, J. 2014. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980.
  18. 18.Lassner, C.; and Zollhofer, M. 2021. Pulsar: Efficient sphere-based neural rendering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 1440–1449.
  19. 19.LeCun, Y.; Bengio, Y.; and Hinton, G. 2015. Deep learning. nature, 521(7553): 436–444.
  20. 20.Li, C.; Feng, B. Y.; Liu, Y.; Liu, H.; Wang, C.; Yu, W.; and Yuan, Y. 2024. Endosparse: Real-time sparse view synthesis of endoscopic scenes using gaussian splatting. In International Conference on Medical Image Computing and Computer-Assisted Intervention, 252–262. Springer.
  21. 21.Li, Z.; Wang, M.; Pi, H.; Xu, K.; Mei, J.; and Liu, Y. 2022. E-nerv: Expedite neural video representation with disentangled spatial-temporal context. In European Conference on Computer Vision, 267–284. Springer.
  22. 22.Liu, X.; Zhan, X.; Tang, J.; Shan, Y.; Zeng, G.; Lin, D.; Liu, X.; and Liu, Z. 2024a. Humangaussian: Text-driven 3d human generation with gaussian splatting. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 6646–6657.
  23. 23.Liu, Y.; Li, C.; Yang, C.; and Yuan, Y. 2024b. Endogaussian: Gaussian splatting for deformable surgical scene reconstruction. arXiv preprint arXiv:2401.12561.
  24. 24.Liu, Z.; Zhu, H.; Zhang, Q.; Fu, J.; Deng, W.; Ma, Z.; Guo, Y.; and Cao, X. 2024c. FINER: Flexible spectral-bias tuning in Implicit NEural Representation by Variable-periodic Activation Functions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2713–2722.
  25. 25.Lu, G.; Zhang, S.; Wang, Z.; Liu, C.; Lu, J.; and Tang, Y. 2024. Manigaussian: Dynamic gaussian splatting for multi-task robotic manipulation. arXiv preprint arXiv:2403.08321.
  26. 26.Ma, C.; Yu, P.; Lu, J.; and Zhou, J. 2022. Recovering realistic details for magnification-arbitrary image super-resolution. IEEE Transactions on Image Processing, 31: 3669–3683.
  27. 27.Martel, J. N.; Lindell, D. B.; Lin, C. Z.; Chan, E. R.; Monteiro, M.; and Wetzstein, G. 2021. Acorn: Adaptive coordinate networks for neural scene representation. arXiv preprint arXiv:2105.02788.
  28. 28.Matsuki, H.; Murai, R.; Kelly, P. H.; and Davison, A. J. 2024. Gaussian splatting slam. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 18039–18048.
  29. 29.Michalkiewicz, M.; Pontes, J. K.; Jack, D.; Baktashmotlagh, M.; and Eriksson, A. 2019. Implicit surface representations as layers in neural networks. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 4743–4752.
  30. 30.Mildenhall, B.; Srinivasan, P. P.; Tancik, M.; Barron, J. T.; Ramamoorthi, R.; and Ng, R. 2021. Nerf: Representing scenes as neural radiance fields for view synthesis. Communications of the ACM, 65(1): 99–106.
  31. 31.Mittermaier, M.; Venkatesh, K. P.; and Kvedar, J. C. 2023. Digital health technology in clinical trials. NPJ Digital Medicine, 6(1): 88.
  32. 32.Muller, T.; Evans, A.; Schied, C.; and Keller, A. 2022. Instant neural graphics primitives with a multiresolution hash encoding. ACM transactions on graphics (TOG), 41(4): 1–15.
  33. 33.Park, J. J.; Florence, P.; Straub, J.; Newcombe, R.; and Lovegrove, S. 2019. Deepsdf: Learning continuous signed distance functions for shape representation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 165–174.
  34. 34.Rabbani, M.; and Joshi, R. 2002. An overview of the JPEG 2000 still image compression standard. Signal processing: Image communication, 17(1): 3–48.
  35. 35.Ramasinghe, S.; and Lucey, S. 2022. Beyond periodicity: Towards a unifying framework for activations in coordinate-mlps. In European Conference on Computer Vision, 142–158. Springer.
  36. 36.Ren, K.; Jiang, L.; Lu, T.; Yu, M.; Xu, L.; Ni, Z.; and Dai, B. 2024. Octree-gs: Towards consistent real-time rendering with lod-structured 3d gaussians. arXiv preprint arXiv:2403.17898.
  37. 37.Saragadam, V.; LeJeune, D.; Tan, J.; Balakrishnan, G.; Veeraraghavan, A.; and Baraniuk, R. G. 2023. Wire: Wavelet implicit neural representations. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 18507–18516.
  38. 38.Saragadam, V.; Tan, J.; Balakrishnan, G.; Baraniuk, R. G.; and Veeraraghavan, A. 2022. Miner: Multiscale implicit neural representation. In European Conference on Computer Vision, 318–333. Springer.
  39. 39.Sitzmann, V.; Martel, J.; Bergman, A.; Lindell, D.; and Wetzstein, G. 2020. Implicit neural representations with periodic activation functions. Advances in neural information processing systems, 33: 7462–7473.
  40. 40.Strumpler, Y.; Postels, J.; Yang, R.; Gool, L. V.; and Tombari, F. 2022. Implicit neural representations for image compression. In European Conference on Computer Vision, 74–91. Springer.
  41. 41.Szymanowicz, S.; Rupprecht, C.; and Vedaldi, A. 2024. Splatter image: Ultra-fast single-view 3d reconstruction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 10208–10217.
  42. 42.Tang, J.; Chen, Z.; Chen, X.; Wang, T.; Zeng, G.; and Liu, Z. 2024. Lgm: Large multi-view gaussian model for high-resolution 3d content creation. arXiv preprint arXiv:2402.05054.
  43. 43.Tang, J.; Ren, J.; Zhou, H.; Liu, Z.; and Zeng, G. 2023. Dreamgaussian: Generative gaussian splatting for efficient 3d content creation. arXiv preprint arXiv:2309.16653.
  44. 44.Wang, Y.; He, X.; Dong, Y.; Lin, Y.; Huang, Y.; and Ding, X. 2024. Cross-Modality Interaction Network for Pansharpening. IEEE Transactions on Geoscience and Remote Sensing.
  45. 45.Xiangli, Y.; Xu, L.; Pan, X.; Zhao, N.; Rao, A.; Theobalt, C.; Dai, B.; and Lin, D. 2022. Bungeenerf: Progressive neural radiance field for extreme multi-scale scene rendering. In European conference on computer vision, 106–122. Springer.
  46. 46.Xu, Q.; Wang, W.; Ceylan, D.; Mech, R.; and Neumann, U. 2019. Disn: Deep implicit surface network for high-quality single-view 3d reconstruction. Advances in neural information processing systems, 32.
  47. 47.Xu, Y.; Shi, Z.; Yifan, W.; Chen, H.; Yang, C.; Peng, S.; Shen, Y.; and Wetzstein, G. 2024. Grm: Large gaussian reconstruction model for efficient 3d reconstruction and generation. arXiv preprint arXiv:2403.14621.
  48. 48.Yan, Z.; Low, W. F.; Chen, Y.; and Lee, G. H. 2024. Multi-scale 3d gaussian splatting for anti-aliased rendering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 20923–20931.
  49. 49.Ye, V.; and Kanazawa, A. 2023. Mathematical Supplement for the gsplat Library. arXiv:2312.02121.
  50. 50.Yu, Z.; Chen, A.; Huang, B.; Sattler, T.; and Geiger, A. 2024. Mip-splatting: Alias-free 3d gaussian splatting. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 19447–19456.
  51. 51.Yugay, V.; Li, Y.; Gevers, T.; and Oswald, M. R. 2023. Gaussian-slam: Photo-realistic dense slam with gaussian splatting. arXiv preprint arXiv:2312.10070.
  52. 52.Zhang, X.; Ge, X.; Xu, T.; He, D.; Wang, Y.; Qin, H.; Lu, G.; Geng, J.; and Zhang, J. 2024a. GaussianImage: 1000 FPS Image Representation and Compression by 2D Gaussian Splatting. arXiv preprint arXiv:2403.08551.
  53. 53.Zhang, Y.; Kuznetsov, A.; Jindal, A.; Chen, K.; Sochenov, A.; Kaplanyan, A.; and Sun, Q. 2024b. Image-GS: Content-Adaptive Image Representation via 2D Gaussians. arXiv preprint arXiv:2407.01866.
  54. 54.Zhao, H.; Zhao, X.; Zhu, L.; Zheng, W.; and Xu, Y. 2024. HFGS: 4D Gaussian Splatting with Emphasis on Spatial and Temporal High-Frequency Components for Endoscopic Scene Reconstruction. arXiv preprint arXiv:2405.17872.
  55. 55.Zhu, L.; Wang, Z.; Jin, Z.; Lin, G.; and Yu, L. 2024. Deformable endoscopic tissues reconstruction with gaussian splatting. arXiv preprint arXiv:2401.11535.

Citation

MLA
Zhu, L., et al. “Large Images Are Gaussians: High-Quality Large Image Representation with Levels of 2D Gaussian Splatting”. arXiv, 2025, http://arxiv.org/abs/2502.09039v1.
APA
Zhu, L., Lin, G., Chen, J., Zhang, X., Jin, Z., Wang, Z., & Yu, L. (2025). Large Images are Gaussians: High-Quality Large Image Representation with Levels of 2D Gaussian Splatting. arXiv. http://arxiv.org/abs/2502.09039v1
Chicago
Zhu, L., G. Lin, J. Chen, et al. 2025. “Large Images Are Gaussians: High-Quality Large Image Representation with Levels of 2D Gaussian Splatting”. arXiv. http://arxiv.org/abs/2502.09039v1.
Harvard
Zhu, L. et al. (2025) “Large Images are Gaussians: High-Quality Large Image Representation with Levels of 2D Gaussian Splatting”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2502.09039v1.
Vancouver
1. Zhu L, Lin G, Chen J, Zhang X, Jin Z, Wang Z, Yu L (2025) Large Images are Gaussians: High-Quality Large Image Representation with Levels of 2D Gaussian Splatting. arXiv

BibTeX

@article{zhu2025large,
  title = {Large Images are Gaussians: High-Quality Large Image Representation with Levels of 2D Gaussian Splatting},
  author = {Zhu, Lingting and Lin, Guying and Chen, Jinnan and Zhang, Xinjie and Jin, Zhenchao and Wang, Zhao and Yu, Lequan},
  year = {2025},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2502.09039v1},
  eprint = {2502.09039}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/