Efficient Hierarchical Entropy Model for Learned Point Cloud Compression

Rui SongChunyang FuShan LiuGe Li

article2023CVPR58 citations

Proposes an efficient octree-based entropy model using hierarchical attention and grouped contexts to achieve linear computational complexity and fast parallel decoding without sacrificing point cloud compression performance.

Listen

Three-dimensional point clouds are essential for modern spatial computing applications, including autonomous driving navigation and immersive virtual environments. However, these datasets typically contain millions of points, creating massive data volumes that strain data storage systems and transmission bandwidth. While recent deep learning methods achieve superior compression rates by using attention mechanisms over large spatial contexts, they suffer from prohibitive computational costs and serial decoding bottlenecks, often requiring several minutes to process a single data frame.

The article develops and evaluates an efficient hierarchical entropy model (EHEM) designed to deliver high-quality point cloud compression while dramatically cutting decoding delays to practical operating speeds. The core objective is to replace computationally heavy global attention and serial dependencies with a scalable, parallel-friendly architecture that preserves global context awareness.

To achieve this, the authors designed a hierarchical attention structure that computes dependencies within localized windows and captures broader spatial context through multi-scale feature downsampling, reducing computational complexity from quadratic to linear relative to context scale. Additionally, they introduced a grouped context framework that splits occupancy sequences into two subsets, enabling parallel decoding across nodes. The approach was validated through rigorous empirical benchmarking against industry baselines and state-of-the-art learned models using standard autonomous vehicle LiDAR datasets, specifically SemanticKITTI and Ford, measuring bitrates, distortion metrics, and runtime latencies.

The findings demonstrate substantial improvements across compression efficiency and operational speed. The proposed model achieves an average bitrate reduction of 19.47% compared to the leading learned baseline (OctAttention) and 28.89% compared to the standard MPEG G-PCC handcrafted codec on the SemanticKITTI benchmark. Crucially, the model slashes decoding latency for high-resolution frames from approximately 708 seconds down to about 3.01 seconds—a reduction of over 99.5%—while a lightweight variant operates in 1.94 seconds. Computational operations scale linearly as context sizes expand, allowing context windows to increase to 8,192 nodes without excessive computational burden.

These results show that neural point cloud compression can achieve practical deployment viability without sacrificing state-of-the-art data reduction capabilities. For decision-makers, this translates to reduced cloud storage footprints and lower bandwidth transmission costs for spatial data workflows, alongside viable execution timelines for downstream applications. The architectural shift demonstrates that group-based parallel decoding effectively bridges the historical trade-off between compression quality and decoding throughput.

Organizations handling large volumes of 3D spatial data should consider transitioning toward hierarchical learned entropy architectures for point cloud storage and transmission pipelines. Deploying teams can evaluate the standard model for maximum bitrate savings or the lightweight configuration where lower latency and memory footprints are prioritized. Further development and pilot testing should focus on optimizing hardware acceleration and validating the model across diverse sensor modalities beyond automotive LiDAR.

Confidence in these findings is high for automotive LiDAR point cloud distributions under standardized evaluation protocols. However, decision-makers should note that the model requires specialized neural network execution environments (such as dedicated graphics processing units) and exhibits moderately higher memory usage and encoding runtimes compared to traditional rule-based codecs.

No sufficiently relevant recommendations were found.

Cover for Efficient Hierarchical Entropy Model for Learned Point Cloud Compression

Abstract

Learning an accurate entropy model is a fundamental way to remove the redundancy in point cloud compression. Recently, the octree-based auto-regressive entropy model which adopts the self-attention mechanism to explore dependencies in a large-scale context is proved to be promising. However, heavy global attention computations and auto-regressive contexts are inefficient for practical applications. To improve the efficiency of the attention model, we propose a hierarchical attention structure that has a linear complexity to the context scale and maintains the global receptive field. Furthermore, we present a grouped context structure to address the serial decoding issue caused by the auto-regression while preserving the compression performance. Experiments demonstrate that the proposed entropy model achieves superior rate-distortion performance and significant decoding latency reduction compared with the state-of-the-art large-scale auto-regressive entropy model.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 2.1. Learned Point Cloud Compression
  • 2.2. Learned Image Compression
  • 3. Preliminary
  • 3.1. Octree Structure
  • 3.2. Large-scale Auto-regressive Entropy Model
  • 4. Efficient Hierarchical Entropy Model
  • 4.1. Overall Architecture
  • 4.2. Grouped Context
  • 4.3. Hierarchical Attention
  • 4.4. Learning
  • 5. Experiments
  • 5.1. Experimental Settings
  • 5.2. Performance Evaluation
  • 5.3. Ablation Studies and Analysis
  • 5.4. Qualitative Results
  • 6. Conclusion
  • Acknowledgement
  • References

Knowls

  1. Knowl 1 — Efficient hierarchical entropy model

    model/method

    EHEM is a learned entropy model for losslessly coding the occupancy-symbol sequence of a quantized point-cloud octree. The octree nodes are traversed breadth-first, and each non-overlapping context window contains NN consecutive occupancy symbols. For every window, EHEM combines two mechanisms: a grouped context that divides the symbols into two decodable groups, and hierarchical attention that models dependencies among ancestral and sibling features.

    For each node, the ancestral context contains the node's known geometric metadata—such as octree level, octant, and parent bounding-box coordinates—together with features from KK known ancestors, while excluding the node's own unknown occupancy symbol. The first symbol group is predicted from ancestral features. After that group is decoded, its symbols are embedded as sibling features and combined with ancestral features to predict the second group. The resulting conditional distributions are passed to an arithmetic coder.

  2. Knowl 2 — Grouped context for parallel decoding

    model/method

    EHEM partitions every context window into two disjoint groups so that symbols within a group have no dependencies on one another. For an even-length window with local symbols z_1,dots,z_N and complete ancestral context AA, the decomposition is

    z(1)={z1,z3,…,zN−1},C(1)={A},z(2)={z2,z4,…,zN},C(2)={A,z(1)}.\begin{aligned} \mathbf{z}^{(1)} &= \{z_1,z_3,\dots,z_{N-1}\}, & \mathcal{C}^{(1)} &= \{A\},\\ \mathbf{z}^{(2)} &= \{z_2,z_4,\dots,z_N\}, & \mathcal{C}^{(2)} &= \{A,\mathbf{z}^{(1)}\}. \end{aligned}

    The first group is encoded or decoded using only ancestral features; the second group additionally uses the already decoded first group. Consequently, all symbols in each group can be processed in parallel, reducing the number of sequential inference steps from NN for a fully autoregressive context to two. The alternating pattern preserves the two nearest decoded neighbors for each even-position symbol while sampling more distant sibling references uniformly.

    Unlike a causal sibling context, the grouped structure can use the complete ancestral context, including ancestral features at positions later than the currently predicted symbol. The authors' attention analysis found that these noncausal ancestral features carry dependency patterns comparable to causal ancestral features, which helps preserve compression performance despite allowing only half of the symbols to access sibling information.

  3. Knowl 3 — Multi-scale hierarchical attention

    model/method

    EHEM replaces global self-attention over the full context with localized attention plus progressive feature merging. The NN ancestral features are divided into non-overlapping local windows of length LL. Parent coordinates within each window are processed by a DGCNN-based geometric feature extractor, producing CC-channel geometry-aware features.

    Localized self-attention is applied independently inside each window. Neighboring tokens are then merged by concatenating their CC-channel features into 2C2C channels and using an MLP to return to CC channels. Successive blocks operate at progressively coarser scales, and alternating shifted-window partitions allow information to cross the boundaries of neighboring windows. A stack of log⁡2(N/L)+1\log_2(N/L)+1 hierarchical blocks eventually gives the representation access to all NN tokens while retaining fine-grained representations for nearby nodes and coarser representations for distant nodes.

    For predicting the second grouped-context subset, EHEM uses hierarchical cross-attention. Each local window contains L/2L/2 first-group sibling features as keys and values and L/2L/2 second-group features as queries. Cross-attention is computed locally, neighboring keys and queries are merged at successive scales, and the multi-scale sibling representation is combined with the ancestral representation before an MLP predicts the second group's occupancy distributions.

  4. Knowl 4 — Linear attention complexity with context scale

    theoretical result

    For an attention layer with NN context tokens, local-window length LL, and feature dimension CC, the computation costs reported for global and localized self-attention are

    Ωglobal=2N2C,Ωlocalized=2LNC.\Omega_{\mathrm{global}}=2N^2C, \qquad \Omega_{\mathrm{localized}}=2LNC.

    Global attention is quadratic in the context length NN, whereas localized attention is linear in NN when LL is fixed. This reduction enables EHEM to use substantially deeper attention stacks and larger context windows without the rapid computation growth of the global autoregressive model. The hierarchical merging and shifted-window operations provide long-range receptive fields while preserving the linear dependence on the number of input tokens for each localized attention layer.

  5. Knowl 5 — Entropy-model training objective

    equation

    Let xix_i be the occupancy symbol of octree node ii, let Ci\mathcal{C}_i be the context available when that node is coded, and let p~i(xi∣Ci)\tilde{p}_i(x_i\mid\mathcal{C}_i) be EHEM's predicted probability for that symbol. EHEM is trained by minimizing the total negative log-likelihood, which corresponds to the estimated arithmetic-coding bitrate when logarithms are base 2:

    ℓ=−∑ilog⁡2p~i(xi∣Ci).\ell=-\sum_i \log_2 \tilde{p}_i(x_i\mid\mathcal{C}_i).

    The sum covers the octree occupancy symbols being communicated. Improving the conditional probability estimates reduces the cross-entropy between the predicted and ground-truth symbol distributions and therefore reduces the coded bitrate without changing the quantized octree geometry.

  6. Knowl 6 — Experimental configuration and model variants

    experimental setup

    EHEM was evaluated on SemanticKITTI and Ford LiDAR point-cloud datasets against the octree-based OctAttention model, the voxel-based SparsePCGC model, and the handcrafted MPEG G-PCC codec. SemanticKITTI contains 43,552 scans; sequences 00–10 were used for training and sequences 11–21 for evaluation. Ford contains three sequences of 1,500 scans; sequence 01 was used for training and sequences 02–03 for evaluation.

    For SemanticKITTI, the quantization step was 400/2D−1400/2^{D-1} for an octree of depth DD, with maximum depth 1616. For Ford, the step was 218−D2^{18-D} with maximum depth 1818; because the original data use 18-bit precision, the highest-depth representation is lossless. Rate was measured in bits per point. Distortion was measured using point-to-point PSNR, point-to-plane PSNR, and Chamfer distance, with PSNR peak values of 59.7059.70 for SemanticKITTI and 3000030000 for Ford.

    The full EHEM used five hierarchical self-attention blocks with 4,4,4,4,24,4,4,4,2 layers, four cross-attention blocks with 2,2,1,12,2,1,1 layers, four attention heads, channel dimension C=256C=256, ancestor depth K=3K=3, context length N=8192N=8192, and local-window length L=512L=512. Light EHEM retained the architecture but used 2,2,2,2,22,2,2,2,2 self-attention layers and C=192C=192. Both variants were optimized with Adam at learning rate 10−410^{-4} for 10 epochs on SemanticKITTI and 50 epochs on Ford, with evaluation on an NVIDIA V100 GPU.

  7. Knowl 7 — Rate-distortion improvement over competing codecs

    empirical result

    Across the SemanticKITTI rate-distortion curves, EHEM achieved a reported average bitrate gain of 28.89%28.89\% relative to G-PCC and a 19.47%19.47\% bitrate reduction relative to OctAttention. EHEM also outperformed the compared methods on Ford across the reported D1 PSNR, D2 PSNR, and Chamfer-distance operating points. Light EHEM had lower computational cost while retaining better rate-distortion performance than the other comparison baselines.

    At approximately matched bitrates on two SemanticKITTI examples, EHEM produced substantially higher D1 PSNR. In one example, G-PCC achieved 80.6280.62 dB at 7.197.19 bits per point, OctAttention achieved 83.2883.28 dB at 7.187.18 bits per point, and EHEM achieved 90.6190.61 dB at 7.187.18 bits per point. In another, G-PCC achieved 80.7680.76 dB at 7.827.82 bits per point, OctAttention achieved 83.0283.02 dB at 7.827.82 bits per point, and EHEM achieved 90.5990.59 dB at 7.787.78 bits per point.

  8. Knowl 8 — Decoding-latency reduction from grouped coding

    data/table

    The following inference times are in seconds for encoding and decoding a SemanticKITTI octree at depths D=12,14,16D=12,14,16. G-PCC values are total codec runtimes. The comparison demonstrates that grouped decoding removes the extreme serial-decoding cost of OctAttention: at depth 16, EHEM decodes in 3.013.01 seconds instead of 708708 seconds, while Light EHEM decodes in 1.941.94 seconds. EHEM encoding is slower than OctAttention because its two groups require two inference stages, but its decoding time is reduced by more than two orders of magnitude.

    Could not parse LaTeX table
  9. Knowl 9 — Measured computation and memory trade-offs

    data/table

    The proposed localized hierarchy trades a larger model and memory footprint than OctAttention for much lower decoding time and better scaling with context length. The first table reports encoding/decoding time, FLOPs, parameter count, and memory for a depth-16 SemanticKITTI octree. FLOPs are measured for one context-window inference; larger windows also predict more nodes in a forward pass.

    Could not parse LaTeX table

    The second table gives FLOPs as the context length NN grows from 512512 to 8192.At8192. At N=8192,EHEMrequires, EHEM requires 184.4GFLOPscomparedwithGFLOPs compared with514.0GFLOPsforOctAttentionandGFLOPs for OctAttention and102.9GFLOPsforLightEHEM.FromGFLOPs for Light EHEM. FromN=4096toto8192,EHEM′sFLOPsincreasebyafactorof, EHEM's FLOPs increase by a factor of 2.26,whereasOctAttention′sincreasebyafactorof, whereas OctAttention's increase by a factor of 3.53$.

    Could not parse LaTeX table
  10. Knowl 10 — Ablation evidence for hierarchy, context scale, and grouping

    empirical result

    Ablation experiments attribute most of EHEM's rate-distortion improvement to its hierarchical attention rather than to grouping alone. With N=1024N=1024 and comparable FLOPs, the hierarchical model used 18 localized self-attention layers, whereas a global-attention replacement used 5 layers; the deeper hierarchical model achieved considerably better rate-distortion performance.

    Increasing the context length from N=1024N=1024 to N=8192N=8192 enlarged the captured spatial region from a limited local area to geometry containing a complete vehicle pattern and a neighboring similar vehicle, and this larger context improved rate-distortion performance. Increasing network depth also improved performance when the remaining architecture was held fixed.

    A grouped-context model and an autoregressive-context model were compared using global attention for a fair comparison. Their rate-distortion performance was comparable, showing that grouping itself did not significantly improve compression accuracy. Its principal benefit was enabling two-step parallel decoding while retaining the complete ancestral context; the major compression gains of full EHEM came from the stronger hierarchical attention blocks.

Coverage note — No substantial contributed material was omitted; background, related work, and proof-free implementation details not needed to reconstruct the method were excluded.

References

  1. 1.Johannes Ballé, Valero Laparra, and Eero P Simoncelli. End-to-end optimized image compression. In International Conference on Learning Representations, 2017.
  2. 2.Johannes Ballé, David Minnen, Saurabh Singh, Sung Jin Hwang, and Nick Johnston. Variational image compression with a scale hyperprior. In International Conference on Learning Representations, 2018.
  3. 3.Jens Behley, Martin Garbade, Andres Milioto, Jan Quenzel, Sven Behnke, Cyrill Stachniss, and Jurgen Gall. Semantickitti: A dataset for semantic scene understanding of lidar sequences. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 9297–9307, 2019.
  4. 4.Sourav Biswas, Jerry Liu, Kelvin Wong, Shenlong Wang, and Raquel Urtasun. Muscle: Multi sweep compression of lidar using deep entropy models. In Advances in Neural Information Processing Systems, pages 22170–22181, 2020.
  5. 5.Zhengxue Cheng, Heming Sun, Masaru Takeuchi, and Jiro Katto. Learned image compression with discretized gaussian mixture likelihoods and attention modules. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 7939–7948, 2020.
  6. 6.MPEG 3D Graphics Coding. Common test conditions for g-pcc. ISO/IEC JTC1/SC29/WG7 N00106, 2021.
  7. 7.MPEG 3D Graphics Coding. Preliminary dataset for ai-based point cloud experiments. ISO/IEC JTC1/SC29/WG7 W21570, 2022.
  8. 8.Haoqi Fan, Bo Xiong, Karttikeya Mangalam, Yanghao Li, Zhicheng Yan, Jitendra Malik, and Christoph Feichtenhofer. Multiscale vision transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 6824–6835, 2021.
  9. 9.Guangchi Fang, Qingyong Hu, Hanyun Wang, Yiling Xu, and Yulan Guo. 3dac: Learning attribute compression for point clouds. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 14819–14828, 2022.
  10. 10.Chunyang Fu, Ge Li, Rui Song, Wei Gao, and Shan Liu. Octattention: Octree-based large-scale contexts model for point cloud compression. In Proceedings of the AAAI Conference on Artificial Intelligence, 2022.
  11. 11.Diogo C Garcia, Tiago A Fonseca, Renan U Ferreira, and Ricardo L de Queiroz. Geometry coding for dynamic voxelized point clouds using octrees and multiple contexts. IEEE Transactions on Image Processing, 29:313–322, 2019.
  12. 12.MPEG Group. Mpeg g-pcc tmc13. https://github.com/MPEGGroup/mpeg-pcc-tmc13, 2021.
  13. 13.Dailan He, Ziming Yang, Weikun Peng, Rui Ma, Hongwei Qin, and Yan Wang. Elic: Efficient learned image compression with unevenly grouped space-channel contextual adaptive coding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5718–5727, 2022.
  14. 14.Dailan He, Yaoyan Zheng, Baocheng Sun, Yan Wang, and Hongwei Qin. Checkerboard context model for efficient learned image compression. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 14771–14780, 2021.
  15. 15.Yun He, Xinlin Ren, Danhang Tang, Yinda Zhang, Xiangyang Xue, and Yanwei Fu. Density-preserving deep point cloud compression. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2333–2342, 2022.
  16. 16.Lila Huang, Shenlong Wang, Kelvin Wong, Jerry Liu, and Raquel Urtasun. Octsqueeze: Octree-structured entropy model for lidar compression. In Proceedings of the IEEE/CVF conference on Computer Vision and Pattern Recognition, pages 1313–1323, 2020.
  17. 17.Tianxin Huang and Yong Liu. 3d point cloud geometry compression on deep learning. In Proceedings of the 27th ACM international conference on multimedia, pages 890–898, 2019.
  18. 18.Emre Can Kaya and Ioan Tabus. Neural network modeling of probabilities for coding the octree representation of point clouds. In 2021 IEEE 23rd International Workshop on Multimedia Signal Processing, 2021.
  19. 19.Jun–Hyuk Kim, Byeongho Heo, and Jong–Seok Lee. Joint global and local hierarchical priors for learned image compression. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5982–5991, 2022.
  20. 20.Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization. In International Conference on Learning Representations, 2015.
  21. 21.Jooyoung Lee, Seunghyun Cho, and Seung-Kwon Beack. Context-adaptive entropy model for end-to-end optimized image compression. In International Conference on Learning Representations, 2019.
  22. 22.Ze Liu, Han Hu, Yutong Lin, Zhuliang Yao, Zhenda Xie, Yixuan Wei, Jia Ning, Yue Cao, Zheng Zhang, Li Dong, et al. Swin transformer v2: Scaling up capacity and resolution. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12009–12019, 2022.
  23. 23.Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 10012–10022, 2021.
  24. 24.Donald Meagher. Geometric modeling using octree encoding. Computer graphics and image processing, 19(2):129–147, 1982.
  25. 25.Rufael Mekuria, Kees Blom, and Pablo Cesar. Design, implementation, and evaluation of a point cloud codec for tele-immersive video. IEEE Transactions on Circuits and Systems for Video Technology, 27(4):828–842, 2016.
  26. 26.Fabian Mentzer, Eirikur Agustsson, Michael Tschannen, Radu Timofte, and Luc Van Gool. Conditional probability models for deep image compression. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 4394–4402, 2018.
  27. 27.Fabian Mentzer, George D Toderici, Michael Tschannen, and Eirikur Agustsson. High-fidelity generative image compression. In Advances in Neural Information Processing Systems, pages 11913–11924, 2020.
  28. 28.David Minnen, Johannes Balle, and George D Toderici. Joint autoregressive and hierarchical priors for learned image compression. Advances in neural information processing systems, 31, 2018.
  29. 29.David Minnen and Saurabh Singh. Channel-wise autoregressive entropy models for learned image compression. In 2020 IEEE International Conference on Image Processing, pages 3339–3343, 2020.
  30. 30.Dat Thanh Nguyen and Andre Kaup. Learning-based lossless point cloud geometry coding using sparse representations. arXiv preprint arXiv:2204.05043, 2022.
  31. 31.Dat Thanh Nguyen, Maurice Quach, Giuseppe Valenzise, and Pierre Duhamel. Lossless coding of point cloud geometry using a deep generative model. IEEE Transactions on Circuits and Systems for Video Technology, 31(12):4617–4629, 2021.
  32. 32.Gaurav Pandey, James R McBride, and Ryan M Eustice. Ford campus vision and lidar data set. The International Journal of Robotics Research, 30(13):1543–1552, 2011.
  33. 33.Charles R Qi, Hao Su, Kaichun Mo, and Leonidas J Guibas. Pointnet: Deep learning on point sets for 3d classification and segmentation. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 652–660, 2017.
  34. 34.Charles R Qi, Li Yi, Hao Su, and Leonidas J Guibas. Pointnet++: Deep hierarchical feature learning on point sets in a metric space. In Advances in Neural Information Processing Systems, 2017.
  35. 35.Yichen Qian, Ming Lin, Xiuyu Sun, Zhiyu Tan, and Rong Jin. Entroformer: A transformer-based entropy model for learned image compression. In International Conference on Learning Representations, 2022.
  36. 36.Maurice Quach, Giuseppe Valenzise, and Frederic Dufaux. Learning convolutional transforms for lossy point cloud geometry compression. In IEEE International Conference on Image Processing, pages 4320–4324, 2019.
  37. 37.Zizheng Que, Guo Lu, and Dong Xu. Voxelcontext-net: An octree based framework for point cloud compression. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6042–6051, 2021.
  38. 38.Ruwen Schnabel and Reinhard Klein. Octree-based point-cloud compression. In Proceedings of the 3rd Eurographics / IEEE VGTC conference on Point-Based Graphics, pages 111–120, 2006.
  39. 39.Sebastian Schwarz, Marius Preda, Vittorio Baroncini, Madhukar Budagavi, Pablo Cesar, Philip A Chou, Robert A Cohen, Maja Krivokuca, Sébastien Lasserre, Zhu Li, et al. Emerging mpeg standards for point cloud compression. IEEE Journal on Emerging and Selected Topics in Circuits and Systems, 9(1):133–148, 2018.
  40. 40.Xihua Sheng, Li Li, Dong Liu, Zhiwei Xiong, Zhu Li, and Feng Wu. Deep-pcac: An end-to-end deep lossy compression framework for point cloud attributes. IEEE Transactions on Multimedia, 24:2617–2632, 2021.
  41. 41.Fei Song, Yiting Shao, Wei Gao, Haiqiang Wang, and Thomas Li. Layer-wise geometry aggregation framework for lossless lidar point cloud compression. IEEE Transactions on Circuits and Systems for Video Technology, 31(12):4603–4616, 2021.
  42. 42.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Advances in Neural Information Processing Systems, page 6000–6010, 2017.
  43. 43.Jianqiang Wang, Dandan Ding, Zhu Li, Xiaoxing Feng, Chuntong Cao, and Zhan Ma. Sparse tensor-based multi-scale representation for point cloud geometry compression. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2022.
  44. 44.Jianqiang Wang, Hao Zhu, Haojie Liu, and Zhan Ma. Lossy point cloud geometry compression via end-to-end learning. IEEE Transactions on Circuits and Systems for Video Technology, 31(12):4909–4923, 2021.
  45. 45.Wenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan, Kaitao Song, Ding Liang, Tong Lu, Ping Luo, and Ling Shao. Pyramid vision transformer: A versatile backbone for dense prediction without convolutions. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 568–578, 2021.
  46. 46.Yue Wang, Yongbin Sun, Ziwei Liu, Sanjay E Sarma, Michael M Bronstein, and Justin M Solomon. Dynamic graph cnn for learning on point clouds. Acm Transactions On Graphics, 38(5):1–12, 2019.
  47. 47.Ian H Witten, Radford M Neal, and John G Cleary. Arithmetic coding for data compression. Communications of the ACM, 30(6):520–540, 1987.
  48. 48.Wei Yan, Shan Liu, Thomas H Li, Zhu Li, Ge Li, et al. Deep autoencoder-based lossy geometry compression for point clouds. arXiv preprint arXiv:1905.03691, 2019.
  49. 49.Yinhao Zhu, Yang Yang, and Taco Cohen. Transformer-based transform coding. In International Conference on Learning Representations, 2022.

Citation

MLA
Song, R., et al. “Efficient Hierarchical Entropy Model for Learned Point Cloud Compression”. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023, pp. 14368–77, https://doi.org/10.1109/CVPR52729.2023.01381.
APA
Song, R., Fu, C., Liu, S., & Li, G. (2023). Efficient Hierarchical Entropy Model for Learned Point Cloud Compression. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 14368–14377. https://doi.org/10.1109/CVPR52729.2023.01381
Chicago
Song, R., C. Fu, S. Liu, and G. Li. 2023. “Efficient Hierarchical Entropy Model for Learned Point Cloud Compression”. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 14368–77. https://doi.org/10.1109/CVPR52729.2023.01381.
Harvard
Song, R. et al. (2023) “Efficient Hierarchical Entropy Model for Learned Point Cloud Compression”, 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp. 14368–14377. Available at: https://doi.org/10.1109/CVPR52729.2023.01381.
Vancouver
1. Song R, Fu C, Liu S, Li G (2023) Efficient Hierarchical Entropy Model for Learned Point Cloud Compression. In: 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp 14368–14377

BibTeX

@inproceedings{Song_2023, title={Efficient Hierarchical Entropy Model for Learned Point Cloud Compression}, url={http://dx.doi.org/10.1109/CVPR52729.2023.01381}, DOI={10.1109/cvpr52729.2023.01381}, booktitle={2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)}, publisher={IEEE}, author={Song, Rui and Fu, Chunyang and Liu, Shan and Li, Ge}, year={2023}, month=June, pages={14368–14377} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE