Towards Large-Scale 3D Representation Learning with Multi-Dataset Point Prompt Training

Xiaoyang WuZhuotao TianXin WenBohao PengXihui LiuKaicheng YuHengshuang Zhao

article2024CVPR113 citations

Proposes Point Prompt Training to overcome negative transfer across diverse 3D point cloud datasets by using domain-specific prompts and language-guided label alignment to train a single high-performing representation model.

Listen

Three-dimensional deep learning has historically lagged behind 2D vision and language processing because 3D point cloud datasets are relatively small and expensive to collect. While combining multiple existing 3D datasets into a single training pool seems like a natural way to scale up models, direct joint training causes severe negative transfer. Differences in data density, sensor environments, and label spaces cause models trained on combined data to perform worse on individual datasets than models trained on a single source alone.

The article aims to overcome these multi-source domain barriers by developing Point Prompt Training, a unified training framework that enables a single 3D perception model to learn collaboratively from multiple datasets without degrading performance across specific tasks.

To achieve this, the authors designed two core mechanisms: Prompt-driven Normalization and Language-guided Categorical Alignment. The normalization module injects lightweight, learnable dataset-specific prompts into the model layers, adjusting feature scaling and shifting to absorb domain-specific variations while keeping shared backbone representations generalizable. The categorical alignment mechanism maps disparate dataset category names into a shared semantic language embedding using a pre-trained text encoder, unifying conflicting label definitions. The framework was evaluated across diverse indoor and outdoor benchmarks—including ScanNet, S3DIS, Structured3D, SemanticKITTI, nuScenes, and Waymo—using both convolutional and transformer backbones under supervised joint training, supervised pre-training, and unsupervised pre-training settings.

The experimental findings show that Point Prompt Training completely reverses negative transfer and consistently outperforms single-dataset and standard joint-training baselines. In indoor semantic segmentation, joint training increased accuracy on ScanNet from 68.9% to 75.7% and on S3DIS from 63.3% to 72.2% Mean Intersection over Union compared to naive joint baselines. In outdoor autonomous driving benchmarks, the unified model improved SemanticKITTI validation accuracy by over 7 percentage points over scratch baselines. Additionally, the representations transferred successfully to instance segmentation and showed high data efficiency, achieving strong performance even when training scenes had as few as 20 annotated points.

These results demonstrate that organizations can train a single, shared-weight 3D model across varied synthetic, real, indoor, and outdoor data sources rather than deploying fragmented, dataset-specific models. This unified approach reduces operational overhead and model maintenance while unlocking the benefits of synthetic-to-real transfer. It establishes that prompt tuning and language grounding can effectively bridge structural domain gaps during representation pre-training.

Organizations developing 3D perception systems should adopt domain prompting and language-based label unification when combining disparate point cloud data sources. Future engineering and research efforts should explore extending the framework to simultaneous multi-task training (such as combining segmentation with 3D object detection) and evaluate joint cross-domain training that mixes indoor and outdoor datasets directly.

Confidence in these findings is high given the extensive testing across standard benchmarks and diverse model architectures. However, decision-makers should note that the current study focuses primarily on dense semantic segmentation tasks, relies on pre-trained text encoders for category alignment, and has not yet evaluated a single model trained simultaneously on mixed indoor-outdoor domain extremes.

No sufficiently relevant recommendations were found.

Cover for Towards Large-Scale 3D Representation Learning with Multi-Dataset Point Prompt Training

Abstract

The rapid advancement of deep learning models is often attributed to their ability to leverage massive training data. In contrast, such privilege has not yet fully benefited 3D deep learning, mainly due to the limited availability of large-scale 3D datasets. Merging multiple available data sources and letting them collaboratively train a single model is a potential solution. However, due to the large domain gap between 3D point cloud datasets, such mixed supervision could adversely affect the model’s performance and lead to degenerated performance (i.e., negative transfer) compared to single-dataset training. In view of this challenge, we introduce Point Prompt Training (PPT), a novel framework for multi-dataset synergistic learning in the context of 3D representation learning that supports multiple pre-training paradigms. Based on this framework, we propose Prompt-driven Normalization, which adapts the model to different datasets with domain-specific prompts and Language-guided Categorical Alignment that decently unifies the multiple-dataset label spaces by leveraging the relationship between label text. Extensive experiments verify that PPT can overcome the negative transfer associated with synergistic learning and produce generalizable representations. Notably, it achieves state-of-the-art performance on each dataset using a single weight-shared model with supervised multi-dataset training. Moreover, when served as a pre-training framework, it outperforms other pre-training approaches regarding representation quality and attains remarkable state-of-the-art performance across over ten diverse downstream tasks spanning both indoor and outdoor 3D scenarios.

Table of Contents

  • 1. Introduction
  • 2. Multi-dataset Synergistic Training
  • 2.1. Problem Setup
  • 2.2. Pilot Study: Uncovering the Negative Transfer
  • 3. Point Prompt Training
  • 3.1. Learning with Domain Prompting
  • 3.2. Categorical Alignment
  • 4. Experiments
  • 4.1. Ablation Study
  • 4.2. Results Comparision
  • 5. Conclusion and Discussion
  • Acknowledgements
  • References

Knowls

  1. Knowl 1 — Point Prompt Training Framework for Multi-dataset 3D Representation Learning

    model/method

    Point Prompt Training (PPT) is a framework designed to perform collaborative representation learning across multiple disparate 3D point cloud datasets without encountering negative transfer. Let nn be the number of datasets, and denote each dataset as Di={(xji,yji)}\mathcal{D}_i = \{(x_j^i, y_j^i)\}, where 1≤i≤n1 \le i \le n, xjix_j^i represents an input 3D point cloud, and yjiy_j^i represents its ground-truth annotations (e.g., element-wise semantic segmentation labels). To adapt a shared backbone network f(⋅;θ)f(\cdot; \theta) parameterized by θ\theta to the unique distributions of each dataset domain, PPT assigns a learnable dd-dimensional domain prompt token ci∈Rdc^i \in \mathbb{R}^d to each dataset Di\mathcal{D}_i, forming the context set C={c1,c2,…,cn}\mathcal{C} = \{c^1, c^2, \dots, c^n\}.

    The multi-dataset joint training objective minimizes the cumulative loss over all datasets by jointly optimizing the backbone parameters θ\theta and the domain prompt vectors C\mathcal{C}:

    arg⁡min⁡θ,C∑i=1n1∣Di∣∑(xji,yji)∈DiL(f(xji,ci;θ),yji)\arg\min_{\theta, \mathcal{C}} \sum_{i=1}^n \frac{1}{|\mathcal{D}_i|} \sum_{(x_j^i, y_j^i) \in \mathcal{D}_i} \mathcal{L}\left(f(x_j^i, c^i; \theta), y_j^i\right)

    where L\mathcal{L} denotes the sample-wise loss function (or an unsupervised loss objective when reformulated for unsupervised pre-training). The domain prompt enables the backbone to model domain-specific variance via specialized adapters while learning generalizable geometric representations across distinct datasets.

  2. Knowl 2 — Prompt-driven Normalization for Domain Context Adaptation

    model/method

    Prompt-driven Normalization (PDNorm) is a domain prompt adapter that incorporates domain-specific contextual priors into 3D neural network backbones by modulating normalization layers. For an intermediate feature representation xx from dataset Di\mathcal{D}_i with associated domain prompt ci∈Rdc^i \in \mathbb{R}^d, PDNorm computes the transformed feature as:

    PDNorm(x,ci)=x−E[xˉ]Var[xˉ]+ϵ⋅γ(ci)+β(ci)\text{PDNorm}(x, c^i) = \frac{x - \mathbb{E}[\bar{x}]}{\sqrt{\text{Var}[\bar{x}] + \epsilon}} \cdot \gamma(c^i) + \beta(c^i)

    where xˉ\bar{x} denotes the feature subset over which statistics are computed depending on the base normalization type (e.g., Batch Normalization or Layer Normalization), ϵ>0\epsilon > 0 is a small constant for numerical stability, and γ(ci)\gamma(c^i) and β(ci)\beta(c^i) are affine scale and shift parameters generated via learnable linear projections of the domain prompt cic^i.

    Key implementation properties of PDNorm include:

    1. Independent Statistics: The running mean E[xˉ]\mathbb{E}[\bar{x}] and running variance Var[xˉ]\text{Var}[\bar{x}] are statisticized and maintained independently for each dataset domain Di\mathcal{D}_i.
    2. Zero-Initialization: The linear projections for γ(ci)\gamma(c^i) and β(ci)\beta(c^i) are initialized to zero at the start of training so that the layer initially behaves as standard normalization without prompt disruption.
    3. Learning Rate Scaling: The base learning rate for prompt-related parameters is scaled by a factor of 0.10.1 relative to the backbone learning rate during training from scratch to prioritize backbone representation learning during early stages.
  3. Knowl 3 — Language-guided Categorical Alignment for Multi-dataset 3D Semantic Supervision

    model/method

    Language-guided Categorical Alignment is a mechanism in Point Prompt Training (PPT) designed to unify inconsistent semantic label spaces across multiple 3D point cloud datasets by projecting point representations into a shared semantic space defined by natural language embeddings.

    For a point representation pkip_k^i extracted from dataset Di\mathcal{D}_i, an MLP projection head maps pkip_k^i to the embedding dimension of a frozen, pre-trained text encoder (e.g., CLIP). The textual category names of all datasets are converted into language embeddings {t1,t2,… }\{t_1, t_2, \dots\} using the text encoder. The model is supervised using an InfoNCE contrastive loss that aligns the projected point feature with the text embedding of its ground-truth class while contrasting it against negative class text embeddings restricted strictly to the category space of dataset Di\mathcal{D}_i:

    Lalign=−log⁡exp⁡(⟨MLP(pki),tyki⟩/τ)∑c∈Yiexp⁡(⟨MLP(pki),tc⟩/τ)\mathcal{L}_{\text{align}} = -\log \frac{\exp(\langle \text{MLP}(p_k^i), t_{y_k^i} \rangle / \tau)}{\sum_{c \in \mathcal{Y}_i} \exp(\langle \text{MLP}(p_k^i), t_c \rangle / \tau)}

    where yki∈Yiy_k^i \in \mathcal{Y}_i is the ground-truth category index for point kk in dataset Di\mathcal{D}_i, Yi\mathcal{Y}_i is the set of valid category indices for dataset Di\mathcal{D}_i, ⟨⋅,⋅⟩\langle \cdot, \cdot \rangle denotes cosine similarity, and τ\tau is a temperature hyperparameter. Restricting negative samples to the dataset-specific label space prevents false-negative penalties between distinct datasets with overlapping or hierarchical semantic definitions.

  4. Knowl 4 — Negative Transfer in Naive Joint Training vs. Point Prompt Training

    data/table

    Directly merging multiple 3D point cloud datasets with large domain gaps (e.g., differences in sensor sparsity, scan complexity, and synthetic vs. real distributions) causes negative transfer in naive multi-dataset training baselines. Point Prompt Training (PPT) overcomes this negative transfer, achieving performance gains across all participating indoor datasets simultaneously using a single shared-weight SparseUNet model.

    Target Data Source Sparsity Complexity Scans Baseline Results (Joint Training Data) Ours (All)
    ScanNet S3DIS Struct.3D All PPT
    ScanNet Real Sparse Large rooms 1613 72.2 71.8 65.9 68.9 (-3.3) 75.7 (+3.5)
    S3DIS Real Dense School office 272 64.1 65.4 62.8 63.3 (-2.1) 72.2 (+6.8)
    Struct.3D Synth. Dense Suite 3500 73.7 74.2 74.5 72.9 (-1.6) 75.8 (+1.3)

    In the naive baseline (columns 6--9), joint training on all three datasets leads to performance drops of 3.3%3.3\% mIoU on ScanNet, 2.1%2.1\% mIoU on S3DIS, and 1.6%1.6\% mIoU on Structured3D compared to training on each dataset individually. Under the same multi-dataset setting, PPT achieves substantial improvements of +3.5%+3.5\% mIoU on ScanNet, +6.8%+6.8\% mIoU on S3DIS, and +1.3%+1.3\% mIoU on Structured3D over single-dataset training.

  5. Knowl 5 — Indoor 3D Semantic Segmentation Benchmark Results

    data/table

    Point Prompt Training (PPT) evaluated across indoor 3D semantic segmentation benchmarks on ScanNet, ScanNet200, and S3DIS Area 5 using SparseUNet and Point Transformer v3 (PTv3) backbones under unsupervised pre-training, supervised joint training, and supervised fine-tuning settings.

    Methods Params ScanNet ScanNet200 S3DIS Area5
    Val mIoU Test mIoU Val mIoU Test mIoU mIoU mAcc
    StratifiedFormer 18.8M 74.3 73.7 - - 72.0 78.1
    PointNeXt 41.6M 71.5 71.2 - - 70.5 77.2
    PTv1 11.4M 70.6 - 27.8 - 70.4 76.5
    PTv2 12.8M 75.4 75.2 30.2 - 71.6 77.9
    SparseUNet 39.2M 72.2 73.6 25.0 25.3 65.4 71.7
    + PC 39.2M 74.1 (+1.9) - 26.2 (+1.2) - 70.3 (+4.9) 76.9 (+5.2)
    + CSC 39.2M 73.8 (+1.6) - 26.4 (+1.4) 24.9 (-0.4) 72.2 (+6.8) -
    + MSC 39.2M 75.5 (+3.3) - 28.8 (+3.8) - 70.1 (+4.7) 77.2 (+5.5)
    + PPT Unsup. (f.t.) 41.0M 75.8 (+3.6) - 30.4 (+5.4) - 71.9 (+6.5) 78.3 (+6.6)
    + PPT Sup. (joint) 41.0M 75.7 (+3.5) 76.6 (+3.0) - - 72.2 (+6.8) 78.0 (+6.3)
    + PPT Sup. (f.t.) 41.0M 76.4 (+4.2) - 31.9 (+6.9) 33.2 (+7.9) 72.7 (+7.3) 78.2 (+6.5)
    PTv3 46.2M 77.5 77.9 35.2 37.8 73.4 77.7
    + PPT Sup. (f.t.) 46.3M 78.6 (+1.1) 78.3 (+0.4) 36.0 (+0.8) 39.3 (+1.5) 74.7 (+1.3) 79.6 (+1.9)

    When integrated with unsupervised pre-training (MSC), PPT improves ScanNet200 validation mIoU from 28.8%28.8\% to 30.4%30.4\% and S3DIS Area 5 mIoU from 70.1%70.1\% to 71.9%71.9\%. Supervised fine-tuning with PTv3 achieves state-of-the-art results: 78.6%78.6\% Val mIoU on ScanNet, 36.0%36.0\% Val mIoU on ScanNet200, and 74.7%74.7\% mIoU on S3DIS Area 5.

  6. Knowl 6 — Outdoor 3D Semantic Segmentation Benchmark Results

    data/table

    Point Prompt Training (PPT) evaluated on large-scale outdoor LiDAR semantic segmentation benchmarks across SemanticKITTI, nuScenes, and Waymo datasets using SparseUNet and Point Transformer v3 (PTv3) backbones.

    Methods Params SemanticKITTI nuScenes Waymo
    Val mIoU Test mIoU Val mIoU Test mIoU Val mIoU Val mAcc
    SPVNAS 10.8M 64.7 66.4 - 77.4 - -
    Cylinder3D 26.1M 64.3 67.8 76.1 77.2 - -
    SphereFormer 32.3M 67.8 74.8 78.4 81.9 69.9 -
    SparseUNet 39.2M 63.8 - 73.3 - 65.9 76.6
    + PPT Sup. (joint) 41.0M 70.9 (+7.1) - 78.5 (+5.2) - 70.0 (+4.1) 79.1 (+2.5)
    + PPT Sup. (f.t.) 41.0M 71.4 (+7.6) - 78.6 (+5.3) - 70.4 (+4.5) 78.9 (+2.3)
    PTv3 46.2M 70.8 74.2 80.4 82.7 71.3 80.5
    + PPT Sup. (f.t.) 46.3M 72.3 (+1.5) 75.5 (+1.3) 81.2 (+0.8) 83.0 (+0.3) 72.1 (+0.8) 81.3 (+0.8)

    Jointly training on all outdoor datasets with a single shared SparseUNet model improves SemanticKITTI validation mIoU by +7.1%+7.1\% (63.8%→70.9%63.8\% \to 70.9\%), nuScenes validation mIoU by +5.2%+5.2\% (73.3%→78.5%73.3\% \to 78.5\%), and Waymo validation mIoU by +4.1%+4.1\% (65.9%→70.0%65.9\% \to 70.0\%). Subsequent dataset-specific fine-tuning on PTv3 yields 75.5%75.5\% test mIoU on SemanticKITTI and 83.0%83.0\% test mIoU on nuScenes.

  7. Knowl 7 — Indoor 3D Instance Segmentation Transfer Performance

    data/table

    Transfer learning performance of Point Prompt Training (PPT) representations fine-tuned on ScanNet and ScanNet200 3D instance segmentation benchmarks using the PointGroup framework.

    Backbone Model Params ScanNet Val ScanNet200 Val
    mAP@25 mAP@50 mAP mAP@25 mAP@50 mAP
    SparseUNet 39.2M 72.8 56.9 36.0 32.2 24.5 15.8
    + PC 39.2M - 58.0 (+1.1) - - 24.9 (+0.4) -
    + CSC 39.2M - 59.4 (+2.5) - - 25.2 (+0.7) -
    + LGround 39.2M - - - - 26.1 (+1.6) -
    + MSC 39.2M 74.7 (+1.9) 59.6 (+2.7) 39.3 (+3.3) 34.3 (+2.1) 26.8 (+2.3) 17.3 (+1.5)
    + PPT (f.t.) 41.0M 76.9 (+4.1) 62.0 (+3.1) 40.7 (+4.7) 36.8 (+4.6) 29.4 (+4.9) 19.4 (+3.6)
    PTv3 46.2M 77.5 61.7 40.9 40.1 33.2 23.1
    + PPT (f.t.) 46.3M 78.9 (+1.4) 63.5 (+1.8) 42.1 (+1.2) 40.8 (+0.7) 34.1 (+0.9) 24.0 (+0.9)

    PPT supervised pre-training with SparseUNet achieves 62.0%62.0\% mAP@50 on ScanNet validation (+2.4%+2.4\% higher than MSC) and 29.4%29.4\% mAP@50 on ScanNet200 validation (+2.6%+2.6\% higher than MSC). With PTv3, PPT reaches 42.1%42.1\% mAP on ScanNet and 24.0%24.0\% mAP on ScanNet200.

  8. Knowl 8 — Data-Efficient 3D Scene Understanding with PPT

    data/table

    Evaluation of unsupervised Point Prompt Training (PPT) on the ScanNet Data Efficient benchmark under limited reconstruction percentages (1%1\% to 100%100\%) and limited annotated points per scene (2020 to Full), using SparseUNet as the backbone.

    Limited Reconstructions Limited Annotations
    Pct. SC CSC MSC PPT Pts. SC CSC MSC PPT
    1% 26.0 28.9 (+2.9) 29.2 (+3.2) 31.3 (+5.3) 20 41.9 55.5 (+13.6) 60.1 (+18.2) 60.6 (+18.7)
    5% 47.8 49.8 (+2.0) 50.7 (+2.9) 52.2 (+4.4) 50 53.9 60.5 (+6.6) 66.8 (+12.9) 67.5 (+13.6)
    10% 56.7 59.4 (+2.7) 61.0 (+4.3) 62.8 (+6.1) 100 62.2 65.9 (+3.7) 69.7 (+7.5) 70.8 (+8.6)
    20% 62.9 64.6 (+1.7) 64.9 (+2.0) 66.4 (+3.5) 200 65.5 68.2 (+2.7) 70.7 (+5.2) 72.2 (+6.7)
    100% 72.2 73.8 (+1.6) 75.3 (+3.1) 75.8 (+3.6) Full 72.2 73.8 (+1.6) 75.3 (+3.1) 75.8 (+3.6)

    PPT consistently outperforms training from scratch (SC) and previous self-supervised methods (CSC and MSC) under all label and data scarcity regimes. At extreme data sparsity (1%1\% reconstructions), PPT attains 31.3%31.3\% mIoU (+5.3%+5.3\% over SC); at extreme label sparsity (2020 annotated points per scene), PPT reaches 60.6%60.6\% mIoU (+18.7%+18.7\% over SC).

  9. Knowl 9 — Ablation of Domain Prompt Adapter Architecture and Training Hyperparameters

    data/table

    Ablations on domain prompt adapter designs, initialization strategies, learning rate scaling, insertion locations, and prompt token dimensions on ScanNet 20-category semantic segmentation using SparseUNet joint training across ScanNet, S3DIS, and Structured3D.

    (a) Prompt Adapter Design (b) Zero-Init (c) LR Scaler
    Metric None Add C.A. PDNorm Init w/o w/ Scaler 1.0 0.1 0.01
    Joint mIoU 68.9 70.9 73.5 75.7 Joint 75.2 75.7 Joint 75.4 75.7 75.2
    F.T. mIoU 73.6 73.8 75.4 76.4 F.T. 75.6 76.4 F.T. 76.0 76.4 75.8
    (d) Prompt Location (e) Prompt Dimension (Length)
    Location Initial Encoder Decoder All Dim 128 256 512 1024
    Joint mIoU 68.7 74.2 73.2 75.7 Joint 75.2 75.7 75.7 75.5
    F.T. mIoU 73.9 74.9 74.7 76.4 F.T. 75.9 76.4 76.1 76.2

    Key empirical conclusions:

    1. Prompt-driven Normalization (PDNorm, 75.7%75.7\% joint mIoU) significantly outperforms Direct Injection (Add, 70.9%70.9\%) and Cross Attention (C.A., 73.5%73.5\%).
    2. Zero-initialization improves fine-tuning mIoU from 75.6%75.6\% to 76.4%76.4\%, and scaling prompt learning rate by 0.10.1 is optimal (76.4%76.4\% vs. 76.0%76.0\% at 1.0).
    3. Inserting prompt adapters across all backbone stages (76.4%76.4\%) outperforms inserting at initial (73.9%73.9\%), encoder-only (74.9%74.9\%), or decoder-only (74.7%74.7\%) stages.
    4. A prompt dimension of d=256d=256 is sufficient (76.4%76.4\% fine-tuning mIoU), with larger dimensions (512512, 10241024) offering no additional benefit.
  10. Knowl 10 — Ablation of Categorical Alignment and Alignment Loss Criteria

    data/table

    Ablation of categorical alignment schemes, loss criteria, and multi-dataset sampling ratios on ScanNet 20-category semantic segmentation evaluated on SparseUNet.

    (a) Categorical Alignment Head (b) Language-Guidance Criteria
    Metric Decoupled Unionized L.G. L.G. w/ tpl. Metric L2 Disc. TC InfoNCE
    Joint mIoU 74.4 75.3 75.7 75.8 Joint mIoU 12.8 65.4 70.2 75.7
    F.T. mIoU 74.7 75.8 76.4 76.0 F.T. mIoU 72.1 73.5 73.2 76.4
    (c) Multi-Dataset Sampling Ratio (Structured3D : ScanNet : S3DIS)
    Dataset 4 : 2 : 1 2 : 2 : 1 2 : 1 : 1 1 : 1 : 1
    ScanNet mIoU 75.7 75.8 73.3 74.6
    S3DIS mIoU 72.2 71.9 71.2 71.9
    Struct.3D mIoU 75.8 73.5 72.7 74.7

    Key takeaways:

    1. Language-guided (L.G.) categorical alignment achieves the highest fine-tuning score (76.4%76.4\% mIoU) compared to Decoupled heads (74.7%74.7\%) and a Unionized shared head (75.8%75.8\%). Adding sentence prompt templates (tpl., e.g., 'A point of [class]') degrades fine-tuning performance slightly to 76.0%76.0\%.
    2. InfoNCE loss is superior to L2 regression (12.8%12.8\% joint due to mode collapse), Discriminative Loss (65.4%65.4\%), and Text-supervised Contrastive (TC) loss (70.2%70.2\%).
    3. A dataset sampling ratio of 4:2:14:2:1 (matching dataset scale/convergence rates) yields optimal multi-dataset balance (75.7%75.7\%, 72.2%72.2\%, 75.8%75.8\% across ScanNet, S3DIS, and Structured3D).
  11. Knowl 11 — Limitations and Scope Boundaries of Point Prompt Training

    limitation

    Point Prompt Training operates within specific architectural and domain boundaries:

    1. Separate Domain Cohorts: Supervised multi-dataset pre-training is performed independently within indoor datasets (ScanNet, S3DIS, Structured3D) or outdoor datasets (SemanticKITTI, nuScenes, Waymo), leaving cross-domain unified pre-training spanning both indoor and outdoor point clouds unaddressed.
    2. Single-Task Formulation: Pre-training is restricted to a single task (scene-level semantic segmentation) rather than a unified multi-task objective (e.g., joint detection, segmentation, and completion).
    3. Prompt Design Scope: While PDNorm demonstrates effectiveness over cross-attention and direct injection, deeper visual-prompting formulations and their synergies with emerging unsupervised pre-training objectives remain open to further exploration.

Coverage note — None was omitted; all key architectural components (PDNorm, Language-guided Categorical Alignment), mathematical formulations, pilot study findings, benchmark evaluations (indoor/outdoor segmentation, instance segmentation, data-efficient learning), detailed ablations, and stated limitations are fully covered.

References

  1. 1.Vamsi Aribandi, Yi Tay, Tal Schuster, Jinfeng Rao, Huaixiu Steven Zheng, Sanket Vaibhav Mehta, Honglei Zhuang, Vinh Q. Tran, Dara Bahri, Jianmo Ni, Jai Gupta, Kai Hui, Sebastian Ruder, and Donald Metzler. Ext5: Towards extreme multi-task scaling for transfer learning. In ICLR, 2022. 1
  2. 2.Iro Armeni, Ozan Sener, Amir R. Zamir, Helen Jiang, Ioannis Brilakis, Martin Fischer, and Silvio Savarese. 3d semantic parsing of large-scale indoor spaces. In CVPR, 2016. 2, 3, 7, 13, 15, 17, 18
  3. 3.Yuki M Asano, Christian Rupprecht, Andrew Zisserman, and Andrea Vedaldi. PASS: An imagenet replacement for self-supervised pretraining without humans. In NeurIPS, 2021. 13
  4. 4.Hyojin Bahng, Ali Jahanian, Swami Sankaranarayanan, and Phillip Isola. Exploring visual prompts for adapting large-scale models. arXiv:2203.17274, 2022. 13
  5. 5.Gilad Baruch, Zhuoyuan Chen, Afshin Dehghan, Tal Dimry, Yuri Feigin, Peter Fu, Thomas Gebauer, Brandon Joffe, Daniel Kurz, Arik Schwartz, and Elad Shulman. ARKitscenes - a diverse real-world dataset for 3d indoor scene understanding using mobile RGB-d data. In NeurIPS Workshops, 2021. 3, 13
  6. 6.Jens Behley, Martin Garbade, Andres Milioto, Jan Quenzel, Sven Behnke, Cyrill Stachniss, and Jurgen Gall. Semantickitti: A dataset for semantic scene understanding of lidar sequences. In ICCV, 2019. 2, 7, 18
  7. 7.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. In NeurIPS, 2020. 3, 13
  8. 8.Holger Caesar, Varun Bankiti, Alex H Lang, Sourabh Vora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Giancarlo Baldan, and Oscar Beijbom. nuscenes: A multi-modal dataset for autonomous driving. In CVPR, 2020. 2, 7, 18
  9. 9.Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. End-to-end object detection with transformers. In ECCV, 2020. 4
  10. 10.Rich Caruana. Multitask learning. Machine learning, 1997. 3
  11. 11.Junbum Cha, Sanghyuk Chun, Kyungjae Lee, Han-Cheol Cho, Seunghyun Park, Yunsung Lee, and Sungrae Park. Swad: Domain generalization by seeking flat minima. In NeurIPS, 2021. 13
  12. 12.Xiaozhi Chen, Huimin Ma, Ji Wan, Bo Li, and Tian Xia. Multi-view 3d object detection network for autonomous driving. In CVPR, 2017. 13
  13. 13.Yukang Chen, Jianhui Liu, Xiangyu Zhang, Xiaojuan Qi, and Jiaya Jia. Largekernel3d: Scaling up kernels in 3d sparse cnns. In CVPR, 2023. 18
  14. 14.Yanbei Chen, Manchen Wang, Abhay Mittal, Zhenlin Xu, Paolo Favaro, Joseph Tighe, and Davide Modolo. Scaledet: A scalable multi-dataset object detector. In CVPR, 2023. 13
  15. 15.Hung-Yueh Chiang, Yen-Liang Lin, Yueh-Cheng Liu, and Winston H Hsu. A unified point-based framework for 3d segmentation. In 3DV, 2019. 18
  16. 16.Christopher Choy, JunYoung Gwak, and Silvio Savarese. 4D spatio-temporal convnets: Minkowski convolutional neural networks. In CVPR, 2019. 3, 7, 8, 13, 16, 17, 18, 19
  17. 17.Pointcept Contributors. Pointcept: A codebase for point cloud perception research. https://github.com/Pointcept/Pointcept, 2023. 7, 8, 16, 18, 19
  18. 18.Spconv Contributors. Spconv: Spatially sparse convolution library. https://github.com/traveller59/spconv, 2022. 3, 16
  19. 19.Ganqu Cui, Shengding Hu, Ning Ding, Longtao Huang, and Zhiyuan Liu. Prototypical verbalizer for prompt-based few-shot tuning. In ACL, 2022. 13
  20. 20.Angela Dai and Matthias Nießner. 3dmv: Joint 3d-multi-view prediction for 3d semantic scene segmentation. In ECCV, 2018. 18
  21. 21.Angela Dai, Angel X. Chang, Manolis Savva, Maciej Halber, Thomas Funkhouser, and Matthias Nießner. ScanNet: Richly-annotated 3d reconstructions of indoor scenes. In CVPR, 2017. 1, 2, 3, 7, 8, 13, 15, 17, 18
  22. 22.Bert De Brabandere, Davy Neven, and Luc Van Gool. Semantic instance segmentation with a discriminative loss function. In CVPR, 2017. 6
  23. 23.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. ImageNet: A large-scale hierarchical image database. In CVPR, 2009. 2, 13
  24. 24.Vincent Dumoulin, Jonathon Shlens, and Manjunath Kudlur. A learned representation for artistic style. In ICLR, 2017. 4
  25. 25.Tianyu Gao, Adam Fisch, and Danqi Chen. Making pre-trained language models better few-shot learners. In ACL, 2021. 13
  26. 26.Chunjiang Ge, Rui Huang, Mixue Xie, Zihang Lai, Shiji Song, Shuang Li, and Gao Huang. Domain adaptation via prompt learning. arXiv:2202.06687, 2022. 13
  27. 27.Priya Goyal, Mathilde Caron, Benjamin Lefaudeux, Min Xu, Pengchao Wang, Vivek Pai, Mannat Singh, Vitaliy Liptchinsky, Ishan Misra, Armand Joulin, et al. Self-supervised pretraining of visual features in the wild. arXiv:2103.01988, 2021. 1, 13
  28. 28.Benjamin Graham, Martin Engelcke, and Laurens van der Maaten. 3d semantic segmentation with submanifold sparse convolutional networks. In CVPR, 2018. 13, 18
  29. 29.Yuxian Gu, Xu Han, Zhiyuan Liu, and Minlie Huang. Ppt: Pre-trained prompt tuning for few-shot learning. In ACL, 2022. 13
  30. 30.Meng-Hao Guo, Jun-Xiong Cai, Zheng-Ning Liu, Tai-Jiang Mu, Ralph R Martin, and Shi-Min Hu. Pct: Point cloud transformer. Computational Visual Media, 2021. 13
  31. 31.Xu Han, Weilin Zhao, Ning Ding, Zhiyuan Liu, and Maosong Sun. Ptr: Prompt tuning with rules for text classification. AI Open, 2022. 13
  32. 32.Kaveh Hassani and Mike Haley. Unsupervised multi-task feature learning on point clouds. In ICCV, 2019. 13
  33. 33.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In CVPR, 2016. 5
  34. 34.Kaiming He, Ross Girshick, and Piotr Dollár. Rethinking imagenet pre-training. In ICCV, 2019. 6
  35. 35.Ji Hou, Benjamin Graham, Matthias Nießner, and Saining Xie. Exploring data-efficient 3d scene understanding with contrastive scene contexts. In CVPR, 2021. 1, 2, 7, 8, 13, 18
  36. 36.Yuenan Hou, Xinge Zhu, Yuexin Ma, Chen Change Loy, and Yikang Li. Point-to-voxel knowledge distillation for lidar semantic segmentation. In CVPR, 2022. 19
  37. 37.Qingyong Hu, Bo Yang, Linhai Xie, Stefano Rosa, Yulan Guo, Zhihua Wang, Niki Trigoni, and Andrew Markham. Randla-net: Efficient semantic segmentation of large-scale point clouds. In CVPR, 2020. 18
  38. 38.Shengding Hu, Ning Ding, Huadong Wang, Zhiyuan Liu, Jingang Wang, Juanzi Li, Wei Wu, and Maosong Sun. Knowledgeable prompt-tuning: Incorporating knowledge into prompt verbalizer for text classification. In ACL, 2022. 13
  39. 39.Zeyu Hu, Mingmin Zhen, Xuyang Bai, Hongbo Fu, and Chiew-lan Tai. Jsenet: Joint semantic segmentation and edge detection network for 3d point clouds. In ECCV, 2020. 18
  40. 40.Xun Huang and Serge Belongie. Arbitrary style transfer in real-time with adaptive instance normalization. In ICCV, 2017. 4
  41. 41.Sergey Ioffe and Christian Szegedy. Batch normalization: Accelerating deep network training by reducing internal covariate shift. In ICML, 2015. 5
  42. 42.Menglin Jia, Luming Tang, Bor-Chun Chen, Claire Cardie, Serge Belongie, Bharath Hariharan, and Ser-Nam Lim. Visual prompt tuning. In ECCV, 2022. 2, 3, 4, 13, 17
  43. 43.Li Jiang, Hengshuang Zhao, Shu Liu, Xiaoyong Shen, Chi-Wing Fu, and Jiaya Jia. Hierarchical point-edge interaction network for point cloud semantic segmentation. In ICCV, 2019. 18
  44. 44.Li Jiang, Hengshuang Zhao, Shaoshuai Shi, Shu Liu, Chi-Wing Fu, and Jiaya Jia. Pointgroup: Dual-set point grouping for 3d instance segmentation. In CVPR, 2020. 8
  45. 45.Chen Ju, Tengda Han, Kunhao Zheng, Ya Zhang, and Weidi Xie. Prompting visual-language models for efficient video understanding. In ECCV, 2022. 2, 13
  46. 46.Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. Scaling laws for neural language models. arXiv:2001.08361, 2020. 1, 13
  47. 47.Dongwan Kim, Yi-Hsuan Tsai, Yumin Suh, Masoud Faraki, Sparsh Garg, Manmohan Chandraker, and Bohyung Han. Learning semantic segmentation from multiple datasets with label shifts. In ECCV, 2022. 2, 13
  48. 48.Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C. Berg, Wan-Yen Lo, Piotr Dollár, and Ross Girshick. Segment anything. In ICCV, 2023. 1
  49. 49.Lingdong Kong, Youquan Liu, Runnan Chen, Yuexin Ma, Xinge Zhu, Yikang Li, Yuenan Hou, Yu Qiao, and Ziwei Liu. Rethinking range view representation for lidar segmentation. In ICCV, 2023. 19
  50. 50.Xin Lai, Jianhui Liu, Li Jiang, Liwei Wang, Hengshuang Zhao, Shu Liu, Xiaojuan Qi, and Jiaya Jia. Stratified transformer for 3d point cloud segmentation. In CVPR, 2022. 7, 18
  51. 51.Xin Lai, Yukang Chen, Fanbin Lu, Jianhui Liu, and Jiaya Jia. Spherical transformer for lidar-based 3d recognition. In CVPR, 2023. 7, 8, 19
  52. 52.Loic Landrieu and Martin Simonovsky. Large-scale point cloud semantic segmentation with superpoint graphs. In CVPR, 2018. 18
  53. 53.Alex H Lang, Sourabh Vora, Holger Caesar, Lubing Zhou, Jiong Yang, and Oscar Beijbom. Pointpillars: Fast encoders for object detection from point clouds. In CVPR, 2019. 13
  54. 54.Huan Lei, Naveed Akhtar, and Ajmal Mian. Seggcn: Efficient 3d point cloud segmentation with fuzzy spherical kernel. In CVPR, 2020. 18
  55. 55.Brian Lester, Rami Al-Rfou, and Noah Constant. The power of scale for parameter-efficient prompt tuning. In EMNLP, 2021. 13
  56. 56.Bo Li, Tianlei Zhang, and Tian Xia. Vehicle detection from 3d lidar using fully convolutional network. In RSS, 2016. 13
  57. 57.Yangyan Li, Rui Bu, Mingchao Sun, Wei Wu, Xinhan Di, and Baoquan Chen. Pointcnn: Convolution on x-transformed points. In NeurIPS, 2018. 18
  58. 58.Haojia Lin, Xiawu Zheng, Lijiang Li, Fei Chao, Shanshan Wang, Yan Wang, Yonghong Tian, and Rongrong Ji. Meta architecture for point cloud analysis. In CVPR, 2023. 18
  59. 59.Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig. Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing. ACM Computing Surveys, 2023. 3, 13
  60. 60.Songtao Liu, Zeming Li, and Jian Sun. Self-emd: Self-supervised object detection without imagenet. arXiv:2011.13677, 2020. 13
  61. 61.Xiao Liu, Yanan Zheng, Zhengxiao Du, Ming Ding, Yujie Qian, Zhilin Yang, and Jie Tang. Gpt understands, too. arXiv:2103.10385, 2021. 13
  62. 62.Youquan Liu, Lingdong Kong, Xiaoyang Wu, Runnan Chen, Xin Li, Liang Pan, Ziwei Liu, and Yuexin Ma. Multi-space alignments towards universal lidar segmentation. In CVPR, 2024. 19
  63. 63.Daniel Maturana and Sebastian Scherer. Voxnet: A 3d convolutional neural network for real-time object recognition. In IROS, 2015. 13
  64. 64.Gaku Narita, Takashi Seno, Tomoya Ishikawa, and Yohsuke Kaji. Panopticfusion: Online volumetric semantic mapping at the level of stuff and things. In IROS, 2019. 18
  65. 65.Aaron van den Oord, Yazhe Li, and Oriol Vinyals. Representation learning with contrastive predictive coding. arXiv:1807.03748, 2018. 5, 15
  66. 66.OpenAI. Gpt-4 technical report. arXiv:2303.08774, 2023. 1
  67. 67.Chunghyun Park, Yoonwoo Jeong, Minsu Cho, and Jaesik Park. Fast point transformer. In CVPR, 2022. 18
  68. 68.William Peebles and Saining Xie. Scalable diffusion models with transformers. arXiv:2212.09748, 2022. 4
  69. 69.Bohao Peng, Xiaoyang Wu, Li Jiang, Yukang Chen, Hengshuang Zhao, Zhuotao Tian, and Jiaya Jia. Oa-cnns: Omni-adaptive sparse cnns for 3d semantic segmentation. In CVPR, 2024. 18, 19
  70. 70.Gilles Puy, Alexandre Boulch, and Renaud Marlet. Using a waffle iron for automotive point cloud semantic segmentation. In ICCV, 2023. 19
  71. 71.Charles R Qi, Hao Su, Kaichun Mo, and Leonidas J Guibas. Pointnet: Deep learning on point sets for 3d classification and segmentation. In CVPR, 2017. 13, 18
  72. 72.Charles R Qi, Li Yi, Hao Su, and Leonidas J Guibas. Pointnet++: Deep hierarchical feature learning on point sets in a metric space. In NeurIPS, 2017. 13, 18
  73. 73.Guocheng Qian, Yuchen Li, Houwen Peng, Jinjie Mai, Hasan Hammoud, Mohamed Elhoseiny, and Bernard Ghanem. Pointnext: Revisiting pointnet++ with improved training and scaling strategies. In NeurIPS, 2022. 7, 18
  74. 74.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In ICML, 2021. 5, 15
  75. 75.René Ranftl, Katrin Lasinger, David Hafner, Konrad Schindler, and Vladlen Koltun. Towards robust monocular depth estimation: Mixing datasets for zero-shot cross-dataset transfer. IEEE TPAMI, 2022. 13
  76. 76.David Rozenberszki, Or Litany, and Angela Dai. Language-grounded indoor 3d semantic segmentation in the wild. In ECCV, 2022. 5, 6, 7, 8, 17
  77. 77.Aditya Sanghi. Info3d: Representation learning on 3d objects using mutual information maximization and contrastive learning. In ECCV, 2020. 13
  78. 78.Jonathan Sauder and Bjarne Sievers. Self-supervised deep learning on point clouds by reconstructing space. In NeurIPS, 2019. 13
  79. 79.Timo Schick and Hinrich Schütze. Exploiting cloze-questions for few-shot text classification and natural language inference. In EACL, 2021. 13
  80. 80.Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut. Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning. In ACL, 2018. 2
  81. 81.Taylor Shin, Yasaman Razeghi, Robert L Logan IV, Eric Wallace, and Sameer Singh. Autoprompt: Eliciting knowledge from language models with automatically generated prompts. In EMNLP, 2020. 13
  82. 82.Nathan Silberman, Derek Hoiem, Pushmeet Kohli, and Rob Fergus. Indoor segmentation and support inference from rgbd images. In ECCV, 2012. 15
  83. 83.Shuran Song, Fisher Yu, Andy Zeng, Angel X Chang, Manolis Savva, and Thomas Funkhouser. Semantic scene completion from a single depth image. In CVPR, 2017. 13
  84. 84.Hang Su, Subhransu Maji, Evangelos Kalogerakis, and Erik G. Learned-Miller. Multi-view convolutional neural networks for 3d shape recognition. In ICCV, 2015. 13
  85. 85.Chen Sun, Abhinav Shrivastava, Saurabh Singh, and Abhinav Gupta. Revisiting unreasonable effectiveness of data in deep learning era. In ICCV, 2017. 13
  86. 86.Pei Sun, Henrik Kretzschmar, Xerxes Dotiwalla, Aurelien Chouard, Vijaysai Patnaik, Paul Tsui, James Guo, Yin Zhou, Yuning Chai, Benjamin Caine, et al. Scalability in perception for autonomous driving: Waymo open dataset. In CVPR, 2020. 2, 7
  87. 87.Haotian Tang, Zhijian Liu, Shengyu Zhao, Yujun Lin, Ji Lin, Hanrui Wang, and Song Han. Searching efficient 3d architectures with sparse point-voxel convolution. In ECCV, 2020. 7, 19
  88. 88.Maxim Tatarchenko, Jaesik Park, Vladlen Koltun, and Qian-Yi Zhou. Tangent convolutions for dense prediction in 3d. In CVPR, 2018. 18
  89. 89.Lyne Tchapmi, Christopher Choy, Iro Armeni, JunYoung Gwak, and Silvio Savarese. Segcloud: Semantic segmentation of 3d point clouds. In 3DV, 2017. 18
  90. 90.Hugues Thomas, Charles R Qi, Jean-Emmanuel Deschaud, Beatriz Marcotegui, François Goulette, and Leonidas J Guibas. Kpconv: Flexible and deformable convolution for point clouds. In ICCV, 2019. 13, 18
  91. 91.Bart Thomee, David A Shamma, Gerald Friedland, Benjamin Elizalde, Karl Ni, Douglas Poland, Damian Borth, and Li-Jia Li. Yfcc100m: The new data in multimedia research. Communications of the ACM, 2016. 13
  92. 92.Yonglong Tian, Olivier J Henaff, and Aaron van den Oord. Divide and contrast: Self-supervised learning from uncurated data. In CVPR, 2021. 1, 13
  93. 93.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothee Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. Llama: Open and efficient foundation language models. arXiv:2302.13971, 2023. 1
  94. 94.Dmitry Ulyanov, Andrea Vedaldi, and Victor Lempitsky. Improved texture networks: Maximizing quality and diversity in feed-forward stylization and texture synthesis. In CVPR, 2017. 4
  95. 95.Simon Vandenhende, Stamatios Georgoulis, Wouter Van Gansbeke, Marc Proesmans, Dengxin Dai, and Luc Van Gool. Multi-task learning for dense prediction tasks: A survey. TPAMI, 2021. 2, 13
  96. 96.Chengyao Wang, Li Jiang, Xiaoyang Wu, Zhuotao Tian, Bohao Peng, Hengshuang Zhao, and Jiaya Jia. Groupcontrast: Semantic-aware self-supervised representation learning for 3d understanding. In CVPR, 2024. 18
  97. 97.Jindong Wang, Cuiling Lan, Chang Liu, Yidong Ouyang, Tao Qin, Wang Lu, Yiqiang Chen, Wenjun Zeng, and Philip Yu. Generalizing to unseen domains: A survey on domain generalization. TPAMI, 2022. 13
  98. 98.Lei Wang, Yuchun Huang, Yaolin Hou, Shenman Zhang, and Jie Shan. Graph attention convolution for point cloud semantic segmentation. In CVPR, 2019. 18
  99. 99.Li Wang, Dong Li, Han Liu, Jinzhang Peng, Lu Tian, and Yi Shan. Cross-dataset collaborative learning for semantic segmentation in autonomous driving. In AAAI, 2022. 2, 13
  100. 100.Peng-Shuai Wang. Octformer: Octree-based transformers for 3D point clouds. In SIGGRAPH, 2023. 18
  101. 101.Shenlong Wang, Simon Suo, Wei-Chiu Ma, Andrei Pokrovsky, and Raquel Urtasun. Deep parametric continuous convolutional neural networks. In CVPR, 2018. 18
  102. 102.Xinlong Wang, Wen Wang, Yue Cao, Chunhua Shen, and Tiejun Huang. Images speak in images: A generalist painter for in-context visual learning. In CVPR, 2023. 1
  103. 103.Yue Wang and Justin M Solomon. Deep closest point: Learning representations for point cloud registration. In ICCV, 2019. 13
  104. 104.Xin Wen, Bingchen Zhao, Anlin Zheng, Xiangyu Zhang, and Xiaojuan Qi. Self-supervised visual representation learning with semantic grouping. In NeurIPS, 2022. 13
  105. 105.Wenxuan Wu, Zhongang Qi, and Li Fuxin. Pointconv: Deep convolutional networks on 3d point clouds. In CVPR, 2019. 18
  106. 106.Wenxuan Wu, Li Fuxin, and Qi Shan. Pointconvformer: Revenge of the point-based convolution. In CVPR, 2023. 18
  107. 107.Xiaoyang Wu, Yixing Lao, Li Jiang, Xihui Liu, and Hengshuang Zhao. Point transformer v2: Grouped vector attention and partition-based pooling. In NeurIPS, 2022. 7, 13, 14, 18, 19
  108. 108.Xiaoyang Wu, Xin Wen, Xihui Liu, and Hengshuang Zhao. Masked scene contrast: A scalable framework for unsupervised 3d representation learning. In CVPR, 2023. 2, 3, 7, 8, 13, 18
  109. 109.Jiahao Xie, Xiaohang Zhan, Ziwei Liu, Yew Soon Ong, and Chen Change Loy. Unsupervised object-level representation learning from scene images. In NeurIPS, 2021. 13
  110. 110.Saining Xie, Jiatao Gu, Demi Guo, Charles R Qi, Leonidas Guibas, and Or Litany. Pointcontrast: Unsupervised pretraining for 3d point cloud understanding. In ECCV, 2020. 1, 2, 7, 8, 13, 18
  111. 111.Zhenda Xie, Yutong Lin, Zheng Zhang, Yue Cao, Stephen Lin, and Han Hu. Propagate yourself: Exploring pixel-level consistency for unsupervised visual representation learning. In CVPR, 2021. 13
  112. 112.Mutian Xu, Runyu Ding, Hengshuang Zhao, and Xiaojuan Qi. Paconv: Position adaptive convolution with dynamic kernel assembling on point clouds. In CVPR, 2021. 18
  113. 113.Xu Yan, Chaoda Zheng, Zhen Li, Sheng Wang, and Shuguang Cui. Pointasnl: Robust point clouds processing using nonlocal neural networks with adaptive sampling. In CVPR, 2020. 18
  114. 114.Xu Yan, Jiantao Gao, Chaoda Zheng, Chao Zheng, Ruimao Zhang, Shuguang Cui, and Zhen Li. 2dpass: 2d priors assisted semantic segmentation on lidar point clouds. In ECCV, 2022. 19
  115. 115.Jiancheng Yang, Qiang Zhang, Bingbing Ni, Linguo Li, Jinxian Liu, Mengdie Zhou, and Qi Tian. Modeling point clouds with self-attention and gumbel subset sampling. In CVPR, 2019. 18
  116. 116.Yu-Qi Yang, Yu-Xiao Guo, Jian-Yu Xiong, Yang Liu, Hao Pan, Peng-Shuai Wang, Xin Tong, and Baining Guo. Swin3d: A pretrained transformer backbone for 3d indoor scene understanding. arXiv:2304.06906, 2023. 3, 15, 18
  117. 117.Lewei Yao, Jianhua Han, Xiaodan Liang, Dan Xu, Wei Zhang, Zhenguo Li, and Hang Xu. Detclipv2: Scalable open-vocabulary object detection pre-training via wordregion alignment. In CVPR, 2023. 2, 13
  118. 118.Yuhang Zang, Wei Li, Kaiyang Zhou, Chen Huang, and Chen Change Loy. Unified vision and language prompt learning. arXiv:2210.07225, 2022. 2, 13
  119. 119.Bo Zhang, Jiakang Yuan, Botian Shi, Tao Chen, Yikang Li, and Yu Qiao. Uni3d: A unified baseline for multi-dataset 3d object detection. In CVPR, 2023. 13
  120. 120.Feihu Zhang, Jin Fang, Benjamin Wah, and Philip Torr. Deep fusionnet for point cloud semantic segmentation. In ECCV, 2020. 18
  121. 121.Hengshuang Zhao, Li Jiang, Chi-Wing Fu, and Jiaya Jia. Pointweb: Enhancing local neighborhood features for point cloud processing. In CVPR, 2019. 13, 18
  122. 122.Hengshuang Zhao, Li Jiang, Jiaya Jia, Philip Torr, and Vladlen Koltun. Point transformer. In ICCV, 2021. 7, 13, 18
  123. 123.Xiangyun Zhao, Samuel Schulter, Gaurav Sharma, Yi-Hsuan Tsai, Manmohan Chandraker, and Ying Wu. Object detection with a unified label space from multiple datasets. In ECCV, 2020. 13
  124. 124.Jia Zheng, Junfei Zhang, Jing Li, Rui Tang, Shenghua Gao, and Zihan Zhou. Structured3d: A large photo-realistic dataset for structured 3d modeling. In ECCV, 2020. 2, 3, 15
  125. 125.Kaiyang Zhou, Jingkang Yang, Chen Change Loy, and Ziwei Liu. Conditional prompt learning for vision-language models. In CVPR, 2022. 3
  126. 126.Kaiyang Zhou, Jingkang Yang, Chen Change Loy, and Ziwei Liu. Learning to prompt for vision-language models. IJCV, 2022. 2, 3, 13
  127. 127.Xingyi Zhou, Vladlen Koltun, and Philipp Krähenbühl. Simple multi-dataset detection. In CVPR, 2022. 2, 13
  128. 128.Xinge Zhu, Hui Zhou, Tai Wang, Fangzhou Hong, Yuexin Ma, Wei Li, Hongsheng Li, and Dahua Lin. Cylindrical and asymmetrical 3d convolution networks for lidar segmentation. In CVPR, 2021. 7, 19

Citation

MLA
Wu, X., et al. “Towards Large-scale 3D Representation Learning with Multi-dataset Point Prompt Training”. arXiv, 2023, http://arxiv.org/abs/2308.09718v2.
APA
Wu, X., Tian, Z., Wen, X., Peng, B., Liu, X., Yu, K., & Zhao, H. (2023). Towards Large-scale 3D Representation Learning with Multi-dataset Point Prompt Training. arXiv. http://arxiv.org/abs/2308.09718v2
Chicago
Wu, X., Z. Tian, X. Wen, et al. 2023. “Towards Large-scale 3D Representation Learning with Multi-dataset Point Prompt Training”. arXiv. http://arxiv.org/abs/2308.09718v2.
Harvard
Wu, X. et al. (2023) “Towards Large-scale 3D Representation Learning with Multi-dataset Point Prompt Training”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2308.09718v2.
Vancouver
1. Wu X, Tian Z, Wen X, Peng B, Liu X, Yu K, Zhao H (2023) Towards Large-scale 3D Representation Learning with Multi-dataset Point Prompt Training. arXiv

BibTeX

@article{wu2023towards,
  title = {Towards Large-scale 3D Representation Learning with Multi-dataset Point Prompt Training},
  author = {Wu, Xiaoyang and Tian, Zhuotao and Wen, Xin and Peng, Bohao and Liu, Xihui and Yu, Kaicheng and Zhao, Hengshuang},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2308.09718v2},
  eprint = {2308.09718}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE