SoftGroup for 3D Instance Segmentation on Point Clouds

Thang VuKookhoi KimTung Minh LuuThanh Xuan NguyenChang D. Yoo

article2022CVPR324 citations

Proposes SoftGroup, a 3D instance segmentation method that associates points with multiple semantic classes during bottom-up grouping to prevent error propagation, achieving substantial accuracy gains and fast inference on ScanNet v2 and S3DIS.

Listen

Three-dimensional instance segmentation—identifying individual objects and their exact boundaries within 3D point cloud data—is a critical computer vision capability for autonomous driving, robotics, and augmented reality. Prevailing approaches rely on a bottom-up process that assigns each point strictly to a single object class before grouping nearby points into instances. However, because local parts of objects can be ambiguous, early classification mistakes propagate directly into the grouping phase. This creates incomplete object boundaries and generates false-positive detections, undermining perception accuracy.

The article demonstrates and evaluates SoftGroup, a two-stage method designed to resolve this error propagation. The primary objective is to improve 3D instance segmentation accuracy by enabling flexible, multi-class point associations followed by targeted object refinement.

The approach integrates a bottom-up grouping stage with a top-down refinement stage. Rather than committing each point to a single label, the bottom-up stage uses a soft probability threshold to allow ambiguous points to associate with multiple candidate classes, creating preliminary point-level proposals. The top-down stage then extracts features from each proposal and processes them through classification, segmentation, and mask-scoring branches to refine valid objects and classify erroneous predictions as background. The authors validated SoftGroup on two standard benchmarks, the ScanNet v2 dataset of indoor scans and the Stanford Large-Scale 3D Indoor Spaces (S3DIS) dataset, measuring precision across varying overlap thresholds.

The evaluation yielded several key findings. First, SoftGroup established a new state of the art, outperforming previous leading methods on the ScanNet v2 hidden test set by 6.2 percentage points in 50% overlap average precision (AP50), reaching 76.1%, and leading in 12 of 18 object categories. Second, on S3DIS Area 5, it outperformed the second-best method by 6.8 percentage points in AP50 (reaching 66.1%) and 8.9 percentage points in overall average precision. Third, ablation experiments demonstrated that the score threshold for soft grouping (optimized at 0.2) and the multi-branch top-down refinement work synergistically, boosting baseline performance by 6.5 percentage points. Finally, the system achieved this performance while maintaining high computational efficiency, processing a full scan in 345 milliseconds on standard hardware.

These findings demonstrate that deferring definitive classification until after proposal generation effectively resolves long-standing error propagation issues without sacrificing execution speed. For real-world systems in robotics and automated navigation, this translates to improved operational safety and reliability through reduced false alarms and better object boundary definition, with negligible latency trade-offs.

Organizations developing 3D perception pipelines should consider adopting two-stage soft-grouping frameworks over strict single-label grouping. As next steps, technical teams can utilize the authors' open-source codebase and pre-trained models to reproduce the results and evaluate performance on domain-specific 3D data. Further testing in outdoor environments and under dynamic conditions is recommended, as the study focuses strictly on benchmark indoor datasets.

Cover for SoftGroup for 3D Instance Segmentation on Point Clouds

Abstract

Existing state-of-the-art 3D instance segmentation methods perform semantic segmentation followed by grouping. The hard predictions are made when performing semantic segmentation such that each point is associated with a single class. However, the errors stemming from hard decision propagate into grouping that results in (1) low overlaps between the predicted instance with the ground truth and (2) substantial false positives. To address the aforementioned problems, this paper proposes a 3D instance segmentation method referred to as SoftGroup by performing bottom-up soft grouping followed by top-down refinement. SoftGroup allows each point to be associated with multiple classes to mitigate the problems stemming from semantic prediction errors and suppresses false positive instances by learning to categorize them as background. Experimental results on different datasets and multiple evaluation metrics demonstrate the efficacy of SoftGroup. Its performance surpasses the strongest prior method by a significant margin of +6.2% on the ScanNet v2 hidden test set and +6.8% on S3DIS Area 5 in terms of AP_50. SoftGroup is also fast, running at 345ms per scan with a single Titan X on ScanNet v2 dataset. The source code and trained models for both datasets are available at \url{this https URL}.

Table of Contents

  • 1 Introduction
  • 2 Related work
  • 3 Method
  • 3.1 Point-wise Prediction Network
  • 3.2 Soft Grouping
  • 3.3 Top-Down Refinement
  • 3.4 Multi-task Learning
  • 4 Experiments
  • 4.1 Experimental Settings
  • 4.2 Benchmarking Results
  • 4.3 Qualitative Analysis
  • 4.4 Ablation Study
  • 5 Conclusion
  • References

Knowls

  1. Knowl 1 — SoftGroup Architecture and Two-Stage Pipeline

    model/method

    SoftGroup is a 3D point cloud instance segmentation framework that addresses the error propagation problem inherent in bottom-up grouping methods that rely on hard one-hot semantic labels. The framework operates in two distinct stages:

    1. Bottom-Up Soft Grouping Stage:

      • A point-wise prediction network takes a point cloud P={pi}i=1NP = \{p_i\}_{i=1}^N (with 3D coordinates and color features) as input, voxelizes it into volumetric grids, and feeds it into a 3D U-Net backbone built with Submanifold Sparse Convolutions.
      • Two parallel two-layer MLP branches predict point-wise continuous semantic scores S={s1,…,sN}∈RN×NclassS = \{s_1, \dots, s_N\} \in \mathbb{R}^{N \times N_{\text{class}}} across NclassN_{\text{class}} classes and offset vectors O={o1,…,oN}∈RN×3O = \{o_1, \dots, o_N\} \in \mathbb{R}^{N \times 3} pointing toward instance centroids.
      • A soft grouping module clusters points into preliminary instance proposals using class-specific score thresholding on SS rather than argmax hard selection, allowing ambiguous points to participate in proposals for multiple candidate semantic classes.
    2. Top-Down Refinement Stage:

      • For each proposal, backbone features belonging to the proposal's points are extracted and processed by a shared lightweight tiny U-Net.
      • Three specialized branches process the refined proposal features: a classification branch predicts proposal object categories and filters false positives into a dedicated background class; a segmentation branch refines binary point masks inside each proposal; and a mask scoring branch predicts the Intersection-over-Union (IoU) of the predicted mask with the ground truth to produce a calibrated confidence score.
  2. Knowl 2 — Soft Grouping Algorithm on Point Clouds

    algorithm

    The soft grouping algorithm generates instance proposals from point-wise continuous semantic probability scores and center offset vectors, preventing hard semantic classification errors from truncating valid object instances or forming erroneous partitions.

    Input: Point set P={p1,…,pN}⊂R3P = \{p_1, \dots, p_N\} \subset \mathbb{R}^3, semantic scores S={s1,…,sN}∈RN×NclassS = \{s_1, \dots, s_N\} \in \mathbb{R}^{N \times N_{\text{class}}}, offset vectors O={o1,…,oN}∈RN×3O = \{o_1, \dots, o_N\} \in \mathbb{R}^{N \times 3}, score threshold τ\tau, grouping bandwidth bb
    Output: Set of candidate instance proposals P\mathcal{P}
    Initialize proposal set P←∅\mathcal{P} \leftarrow \emptyset
    for each point i∈{1,…,N}i \in \{1, \dots, N\} do
        p~i←pi+oi\tilde{p}_i \leftarrow p_i + o_i // Shift point toward estimated instance center
    end for
    for each semantic class index c∈{1,…,Nclass}c \in \{1, \dots, N_{\text{class}}\} do
        Vc←{i∈{1,…,N}∣Si,c>τ}V_c \leftarrow \{i \in \{1, \dots, N\} \mid S_{i, c} > \tau\} // Slice point subset exceeding threshold
        if VcV_c is not empty then
            // Construct connected components based on spatial proximity
            Partition VcV_c into disjoint clusters {C1,…,Cm}\{C_1, \dots, C_m\} such that two shifted points p~u,p~v\tilde{p}_u, \tilde{p}_v belong to the same cluster if there exists a path between them with pairwise distance <b< b
            for each cluster CkC_k in {C1,…,Cm}\{C_1, \dots, C_m\} do
                P←P∪{Ck}\mathcal{P} \leftarrow \mathcal{P} \cup \{C_k\}
            end for
        end if
    end for
    return P\mathcal{P}

    Because points can satisfy Si,c>τS_{i, c} > \tau for multiple classes simultaneously, points in locally ambiguous regions are evaluated across multiple candidate classes. The top-down refinement stage subsequently evaluates and prunes false-positive proposals.

  3. Knowl 3 — Top-Down Refinement Network for Instance Proposal Classification and Scoring

    model/method

    The top-down refinement stage processes each candidate instance proposal generated by the soft grouping stage to refine the instance geometry, predict its class, and estimate mask quality:

    • Feature Extraction: A per-instance feature extractor extracts the point-level features directly from the 3D U-Net backbone for the point subset comprising proposal kk. These features are processed by a tiny U-Net.

    • Classification Branch: A global average pooling layer aggregates all point features within proposal kk, followed by an MLP that outputs logits ck∈RNclass+1c_k \in \mathbb{R}^{N_{\text{class}} + 1} over NclassN_{\text{class}} foreground semantic classes and one background class. Classifying the entire proposal suppresses false positives generated by noisy semantic regions by assigning them to the background category.

    • Segmentation Branch: A point-wise two-layer MLP predicts a binary mask mkm_k over the proposal's points, segmenting foreground object points from extraneous background points included during soft grouping.

    • Mask Scoring Branch: Uses the globally pooled proposal features to predict an IoU score vector ek∈RNclasse_k \in \mathbb{R}^{N_{\text{class}}}, estimating the mask overlap with the ground truth.

    • Final Confidence Score: The final confidence score for proposal kk is calculated as the product of the classification confidence score and the mask IoU score: sfinal=ck(c^)⋅ek(c^)s_{\text{final}} = c_k(\hat{c}) \cdot e_k(\hat{c}) where c^\hat{c} is the predicted foreground class.

  4. Knowl 4 — Multi-Task Learning Objective and Target Assignment Formulation

    equation

    SoftGroup is trained end-to-end using a joint multi-task loss function:

    L=Lsemantic+Loffset+Lclass+Lmask+Lmask_scoreL = L_{\text{semantic}} + L_{\text{offset}} + L_{\text{class}} + L_{\text{mask}} + L_{\text{mask\_score}}

    The individual component losses are defined as follows:

    1. Semantic Loss: Lsemantic=1N∑i=1NCE(si,si∗)L_{\text{semantic}} = \frac{1}{N} \sum_{i=1}^N \text{CE}(s_i, s_i^*) where si∈RNclasss_i \in \mathbb{R}^{N_{\text{class}}} is the predicted semantic probability vector for point pip_i, si∗s_i^* is the ground-truth semantic class label, and CE\text{CE} denotes multi-class cross-entropy.

    2. Offset Loss: Loffset=1∑i=1NI{pi}∑i=1NI{pi}∥oi−oi∗∥1L_{\text{offset}} = \frac{1}{\sum_{i=1}^N \mathbb{I}_{\{p_i\}}} \sum_{i=1}^N \mathbb{I}_{\{p_i\}} \|o_i - o_i^*\|_1 where oi∈R3o_i \in \mathbb{R}^3 is the predicted center offset vector, oi∗∈R3o_i^* \in \mathbb{R}^3 is the displacement from point pip_i to the centroid of its ground-truth instance, and I{pi}\mathbb{I}_{\{p_i\}} is an indicator function equal to 1 if point pip_i belongs to an instance and 0 otherwise.

    3. Classification Loss: Lclass=1K∑k=1KCE(ck,ck∗)L_{\text{class}} = \frac{1}{K} \sum_{k=1}^K \text{CE}(c_k, c_k^*) where ck∈RNclass+1c_k \in \mathbb{R}^{N_{\text{class}}+1} is the predicted category logit for proposal kk, KK is the total number of proposals, and ck∗c_k^* is the assigned target label.

    4. Mask Segmentation Loss: Lmask=1∑k=1KI{mk}∑k=1KI{mk}BCE(mk,mk∗)L_{\text{mask}} = \frac{1}{\sum_{k=1}^K \mathbb{I}_{\{m_k\}}} \sum_{k=1}^K \mathbb{I}_{\{m_k\}} \text{BCE}(m_k, m_k^*) where mkm_k is the predicted point mask for proposal kk, mk∗m_k^* is the binary ground-truth instance mask, BCE\text{BCE} is binary cross-entropy, and I{mk}\mathbb{I}_{\{m_k\}} equals 1 if proposal kk is a positive sample.

    5. Mask Scoring Loss: Lmask_score=1∑k=1KI{ek}∑k=1KI{ek}∥ek−ek∗∥22L_{\text{mask\_score}} = \frac{1}{\sum_{k=1}^K \mathbb{I}_{\{e_k\}}} \sum_{k=1}^K \mathbb{I}_{\{e_k\}} \|e_k - e_k^*\|_2^2 where eke_k is the predicted IoU score, ek∗e_k^* is the actual IoU between predicted mask mkm_k and the matched ground-truth mask, and I{ek}\mathbb{I}_{\{e_k\}} equals 1 for positive proposals.

    Sample Assignment Target: An instance proposal is designated as a positive sample (I=1\mathbb{I} = 1) if its IoU with any ground-truth instance exceeds 0.5, assigned to the ground-truth instance with the highest IoU. Proposals with maximum IoU≤0.5\text{IoU} \le 0.5 are treated as negative samples and assigned to the background category (ck∗=backgroundc_k^* = \text{background}) for LclassL_{\text{class}}, while being omitted from LmaskL_{\text{mask}} and Lmask_scoreL_{\text{mask\_score}}.

  5. Knowl 5 — ScanNet v2 Hidden Test Set Instance Segmentation Benchmark

    data/table

    SoftGroup was evaluated on the 3D instance segmentation hidden test split of ScanNet v2 across 18 object classes. The evaluation metric is Average Precision at 50% IoU (AP50\text{AP}_{50}):

    Method AP50_{50} bath bed bksh cab chair cntr curt desk door oth pic frid scur sink sofa tabl toil wind
    SGPN 14.3 20.8 39.0 16.9 6.5 27.5 2.9 6.9 0.0 8.7 4.3 1.4 2.7 0.0 11.2 35.1 16.8 43.8 13.8
    GSPN 30.6 50.0 40.5 31.1 34.8 58.9 5.4 6.8 12.6 28.3 29.0 2.8 21.9 21.4 33.1 39.6 27.5 82.1 24.5
    3D-SIS 38.2 100.0 43.2 24.5 19.0 57.7 1.3 26.3 3.3 32.0 24.0 7.5 42.2 85.7 11.7 69.9 27.1 88.3 23.5
    MASC 44.7 52.8 55.5 38.1 38.2 63.3 0.2 50.9 26.0 36.1 43.2 32.7 45.1 57.1 36.7 63.9 38.6 98.0 27.6
    PanopticFusion 47.8 66.7 71.2 59.5 25.9 55.0 0.0 61.3 17.5 25.0 43.4 43.7 41.1 85.7 48.5 59.1 26.7 94.4 35.9
    3D-BoNet 48.8 100.0 67.2 59.0 30.1 48.4 9.8 62.0 30.6 34.1 25.9 12.5 43.4 79.6 40.2 49.9 51.3 90.9 43.9
    MTML 54.9 100.0 80.7 58.8 32.7 64.7 0.4 81.5 18.0 41.8 36.4 18.2 44.5 100.0 44.2 68.8 57.1 100.0 39.6
    3D-MPA 61.1 100.0 83.3 76.5 52.6 75.6 13.6 58.8 47.0 43.8 43.2 35.8 65.0 85.7 42.9 76.5 55.7 100.0 43.0
    Dyco3D 64.1 100.0 84.1 89.3 53.1 80.2 11.5 58.8 44.8 43.8 53.7 43.0 55.0 85.7 53.4 76.4 65.7 98.7 56.8
    PE 64.5 100.0 77.3 79.8 53.8 78.6 8.8 79.9 35.0 43.5 54.7 54.5 64.6 93.3 56.2 76.1 55.6 99.7 50.1
    PointGroup 63.6 100.0 76.5 62.4 50.5 79.7 11.6 69.6 38.4 44.1 55.9 47.6 59.6 100.0 66.6 75.6 55.6 99.7 51.3
    GICN 63.8 100.0 89.5 80.0 48.0 67.6 14.4 73.7 35.4 44.7 40.0 36.5 70.0 100.0 56.9 83.6 59.9 100.0 47.3
    OccuSeg 67.2 100.0 75.8 68.2 57.6 84.2 47.7 50.4 52.4 56.7 58.5 45.1 55.7 100.0 75.1 79.7 56.3 100.0 46.7
    SSTNet 69.8 100.0 69.7 88.8 55.6 80.3 38.7 62.6 41.7 55.6 58.5 70.2 60.0 100.0 82.4 72.0 69.2 100.0 50.9
    HAIS 69.9 100.0 84.9 82.0 67.5 80.8 27.9 75.7 46.5 51.7 59.6 55.9 60.0 100.0 65.4 76.7 67.6 99.4 56.0
    SoftGroup 76.1 100.0 80.8 84.5 71.6 86.2 24.3 82.4 65.5 62.0 73.4 69.9 79.1 98.1 71.6 84.4 76.9 100.0 59.4

    SoftGroup achieved an average AP50\text{AP}_{50} of 76.1%, outperforming the prior leading method (HAIS at 69.9%) by +6.2% and achieving the best individual score in 12 of the 18 evaluated classes.

  6. Knowl 6 — S3DIS 3D Instance Segmentation Benchmark Results

    data/table

    SoftGroup performance was benchmarked on the Stanford Large-Scale 3D Indoor Spaces (S3DIS) dataset across 13 classes under two evaluation setups: testing on Area 5 and 6-fold cross-validation. Metrics reported are mean Average Precision (AP), Average Precision at 50% IoU (AP50\text{AP}_{50}), mean coverage (mCov), mean weighted coverage (mWCov), mean precision at 50% IoU (mPrec50\text{mPrec}_{50}), and mean recall at 50% IoU (mRec50\text{mRec}_{50}):

    Method AP AP50_{50} mCov mWCov mPrec50_{50} mRec50_{50}
    Evaluated on Area 5
    SGPN - - 32.7 35.5 36.0 28.7
    ASIS - - 44.6 47.8 55.3 42.4
    PointGroup - 57.8 - - 61.9 62.1
    SSTNet 42.7 59.3 - - 65.5 64.2
    HAIS - - 64.3 66.0 71.1 65.0
    SoftGroup 51.6 66.1 66.1 68.0 73.6 66.6
    Evaluated on 6-fold cross validation
    SGPN - - 37.9 40.8 38.2 31.2
    PartNet - - - - 56.4 43.4
    ASIS - - 51.2 55.1 63.6 47.5
    3D-BoNet - - - - 65.6 47.7
    OccuSeg - - - - 72.8 60.3
    GICN - - - - 68.5 50.8
    PointGroup - 64.0 - - 69.6 69.2
    SSTNet 54.1 67.8 - - 73.5 73.4
    HAIS - - 67.0 70.4 73.2 69.4
    SoftGroup 54.4 68.9 69.3 71.7 75.3 69.8

    On Area 5, SoftGroup improves AP and AP50\text{AP}_{50} by +8.9% and +6.8% over the second-best methods, respectively.

  7. Knowl 7 — ScanNet v2 Validation Mask and 3D Bounding Box Object Detection Performance

    data/table

    Performance of 3D instance mask prediction and 3D axis-aligned bounding box prediction extracted from the predicted instance masks on the ScanNet v2 validation set:

    Method AP50_{50} AP25_{25} Box AP50_{50} Box AP25_{25}
    F-PointNet - - 10.8 19.8
    GSPN 37.8 53.4 17.7 30.6
    3D-SIS 18.7 35.7 22.5 40.2
    VoteNet - - 33.5 58.6
    3D-MPA 51.9 72.4 49.2 64.2
    PointGroup 51.7 71.3 48.9 61.5
    SSTNet 64.3 74.0 52.7 62.5
    HAIS 64.4 75.6 53.1 64.3
    SoftGroup 67.6 78.9 59.4 71.6

    Extracting tight axis-aligned bounding boxes from the point masks generated by SoftGroup yields gains over prior state-of-the-art of +3.2% in AP50\text{AP}_{50}, +3.3% in AP25\text{AP}_{25}, +6.3% in Box AP50\text{Box AP}_{50}, and +7.3% in Box AP25\text{Box AP}_{25}.

  8. Knowl 8 — Inference Latency and Component-Wise Execution Time

    data/table

    Inference latency per 3D scene scan on the ScanNet v2 validation set measured on a single NVIDIA Titan X GPU:

    Method Component Breakdown Total Runtime (ms)
    SGPN Backbone (GPU): 2080 ms, Group merging (CPU): 149000 ms, Block merging (CPU): 7119 ms 158439
    ASIS Backbone (GPU): 2083 ms, Mean shift (CPU): 172711 ms, Block merging (CPU): 7119 ms 181913
    GSPN Backbone (GPU): 1612 ms, Point sampling (GPU): 9559 ms, Neighbour search (CPU): 1500 ms 12702
    3D-BoNet Backbone (GPU): 2083 ms, SCN (GPU): 667 ms, Block merging (CPU): 7119 ms 9202
    GICN Backbone (GPU): 1497 ms, SCN (GPU): 667 ms, Block merging (CPU): 7119 ms 8615
    OccuSeg Backbone (GPU): 189 ms, Supervoxel (CPU): 1202 ms, Clustering (GPU+CPU): 513 ms 1904
    PointGroup Backbone (GPU): 128 ms, Clustering (GPU+CPU): 221 ms, ScoreNet (GPU): 103 ms 452
    SSTNet Backbone (GPU): 125 ms, Tree network (GPU+CPU): 229 ms, ScoreNet (GPU): 74 ms 428
    HAIS Pointwise prediction (GPU): 154 ms, Hier. aggr. (GPU+CPU): 118 ms, Intra-inst. (GPU): 67 ms 339
    SoftGroup Pointwise prediction (GPU): 152 ms, Soft grouping (GPU+CPU): 123 ms, Top-down refinement (GPU): 70 ms 345

    SoftGroup achieves 345 ms total inference latency per scene scan, running within 6 ms of the fastest prior model (HAIS at 339 ms) while providing higher segmentation accuracy.

  9. Knowl 9 — Component-Wise Ablation and Threshold Sensitivity in SoftGroup

    empirical result

    Ablation experiments conducted on the ScanNet v2 validation set demonstrate the individual and combined impact of SoftGroup's design choices:

    1. Component-Wise Ablation:

      • Hard grouping baseline with ScoreNet ranking: AP=39.5%\text{AP} = 39.5\%, AP50=61.1%\text{AP}_{50} = 61.1\%, AP25=75.5%\text{AP}_{25} = 75.5\%.
      • Adding Soft Grouping only: AP=41.6%\text{AP} = 41.6\%, AP50=63.8%\text{AP}_{50} = 63.8\%, AP25=79.2%\text{AP}_{25} = 79.2\% (+2.1% AP, +2.7% AP50\text{AP}_{50}).
      • Adding Top-Down Refinement only: AP=44.3%\text{AP} = 44.3\%, AP50=65.4%\text{AP}_{50} = 65.4\%, AP25=78.1%\text{AP}_{25} = 78.1\% (+4.8% AP, +4.3% AP50\text{AP}_{50}).
      • Full SoftGroup (Soft Grouping + Top-Down Refinement): AP=46.0%\text{AP} = 46.0\%, AP50=67.6%\text{AP}_{50} = 67.6\%, AP25=78.9%\text{AP}_{25} = 78.9\% (+6.5% AP, +6.5% AP50\text{AP}_{50}, +3.4% AP25\text{AP}_{25} over baseline).
    2. Soft Grouping Score Threshold (τ\tau):

      • Without threshold (hard grouping baseline): AP=44.3%\text{AP} = 44.3\%, AP50=65.4%\text{AP}_{50} = 65.4\%.
      • τ=0.01\tau = 0.01: AP=40.1%\text{AP} = 40.1\%, AP50=58.5%\text{AP}_{50} = 58.5\%
      • τ=0.1\tau = 0.1: AP=45.3%\text{AP} = 45.3\%, AP50=66.5%\text{AP}_{50} = 66.5\%
      • τ=0.2\tau = 0.2: AP=46.0%\text{AP} = 46.0\%, AP50=67.6%\text{AP}_{50} = 67.6\%, AP25=78.9%\text{AP}_{25} = 78.9\% (optimal balance between foreground recall and precision around 50% precision).
      • τ=0.3\tau = 0.3: AP=45.2%\text{AP} = 45.2\%, AP50=66.8%\text{AP}_{50} = 66.8\%
      • τ=0.4\tau = 0.4: AP=44.7%\text{AP} = 44.7\%, AP50=46.1%\text{AP}_{50} = 46.1\%
      • τ=0.5\tau = 0.5: AP=43.9%\text{AP} = 43.9\%, AP50=64.8%\text{AP}_{50} = 64.8\%
    3. Instance Category Assignment Source:

      • Majority voting of point-wise semantic predictions: AP=45.0%\text{AP} = 45.0\%, AP50=65.6%\text{AP}_{50} = 65.6\%, AP25=76.2%\text{AP}_{25} = 76.2\%.
      • Classification branch logits on pooled proposal features: AP=46.0%\text{AP} = 46.0\%, AP50=67.6%\text{AP}_{50} = 67.6\%, AP25=78.9%\text{AP}_{25} = 78.9\% (+1.0% AP, +2.0% AP50\text{AP}_{50}).
  10. Knowl 10 — SoftGroup Implementation Details and Training Setup

    experimental setup

    The training and inference hyperparameters for SoftGroup are configured as follows:

    • Voxelization and Bandwidth: Point cloud scenes are voxelized with a grid size of 0.02 m0.02\,\text{m}. The grouping bandwidth bb for spatial clustering is set to 0.04 m0.04\,\text{m}.
    • Soft Grouping Score Threshold: Score threshold τ\tau is set to 0.20.2.
    • Optimization and Schedule: The network is trained for 120k iterations using the Adam optimizer with a batch size of 4. The initial learning rate is 0.0010.001, decayed using a cosine annealing schedule.
    • Data Handling: During training on ScanNet v2, scenes are randomly cropped to a maximum limit of 250k250\text{k} points. At inference, full uncropped scenes are processed.
    • S3DIS High-Density Adaptation: Due to higher point density on S3DIS, scenes are randomly downsampled at a ratio of 1/41/4 prior to cropping. At inference, each scene is split into 4 spatial partitions, fed through the backbone, and the extracted point feature maps are merged before proposal generation and top-down refinement.

Coverage note — No substantial contributed material was omitted.

References

  1. 1.Iro Armeni, Ozan Sener, Amir R Zamir, Helen Jiang, Ioannis Brilakis, Martin Fischer, and Silvio Savarese. 3d semantic parsing of large-scale indoor spaces. In CVPR, 2016.
  2. 2.Mathieu Aubry, Ulrich Schlickewei, and Daniel Cremers. The wave kernel signature: A quantum mechanical approach to shape analysis. In ICCV workshops, 2011.
  3. 3.Michael M Bronstein and Iasonas Kokkinos. Scale-invariant heat kernel signatures for non-rigid shape recognition. In CVPR, 2010.
  4. 4.Shaoyu Chen, Jiemin Fang, Qian Zhang, Wenyu Liu, and Xinggang Wang. Hierarchical aggregation for 3d instance segmentation. In ICCV, 2021.
  5. 5.Christopher Choy, JunYoung Gwak, and Silvio Savarese. 4d spatio-temporal convnets: Minkowski convolutional neural networks. In CVPR, 2019.
  6. 6.Angela Dai, Angel X Chang, Manolis Savva, Maciej Halber, Thomas Funkhouser, and Matthias Nießner. Scannet: Richly-annotated 3d reconstructions of indoor scenes. In CVPR, 2017.
  7. 7.Francis Engelmann, Martin Bokeloh, Alireza Fathi, Bastian Leibe, and Matthias Nießner. 3d-mpa: Multi-proposal aggregation for 3d semantic instance segmentation. In CVPR, 2020.
  8. 8.Benjamin Graham, Martin Engelcke, and Laurens Van Der Maaten. 3d semantic segmentation with submanifold sparse convolutional networks. In CVPR, 2018.
  9. 9.Lei Han, Tian Zheng, Lan Xu, and Lu Fang. Occuseg: Occupancy-aware 3d instance segmentation. In CVPR, 2020.
  10. 10.Kaiming He, Georgia Gkioxari, Piotr Dollár, and Ross Girshick. Mask r-cnn. In ICCV, 2017.
  11. 11.Tong He, Chunhua Shen, and Anton van den Hengel. Dyco3d: Robust instance segmentation of 3d point clouds through dynamic convolution. In CVPR, 2021.
  12. 12.Ji Hou, Angela Dai, and Matthias Nießner. 3d-sis: 3d semantic instance segmentation of rgb-d scans. In CVPR, 2019.
  13. 13.Binh-Son Hua, Minh-Khoi Tran, and Sai-Kit Yeung. Pointwise convolutional neural networks. In CVPR, 2018.
  14. 14.Zhaojin Huang, Lichao Huang, Yongchao Gong, Chang Huang, and Xinggang Wang. Mask scoring r-cnn. In CVPR, 2019.
  15. 15.Li Jiang, Hengshuang Zhao, Shaoshuai Shi, Shu Liu, Chi-Wing Fu, and Jiaya Jia. Pointgroup: Dual-set point grouping for 3d instance segmentation. In CVPR, 2020.
  16. 16.Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization. In ICVLR, 2015.
  17. 17.Jean Lahoud, Bernard Ghanem, Marc Pollefeys, and Martin R Oswald. 3d instance segmentation via multi-task metric learning. In ICCV, 2019.
  18. 18.Juho Lee, Yoonho Lee, Jungtaek Kim, Adam Kosiorek, Seungjin Choi, and Yee Whye Teh. Set transformer: A framework for attention-based permutation-invariant neural networks. In ICML, 2019.
  19. 19.Yangyan Li, Rui Bu, Mingchao Sun, Wei Wu, Xinhan Di, and Baoquan Chen. Pointcnn: Convolution on x-transformed points. In NIPS, 2018.
  20. 20.Zhihao Liang, Zhihao Li, Songcen Xu, Mingkui Tan, and Kui Jia. Instance segmentation in 3d scenes using semantic superpoint tree networks. In ICCV, 2021.
  21. 21.Chen Liu and Yasutaka Furukawa. Masc: Multi-scale affinity with sparse convolution for 3d instance segmentation. arXiv:1902.04478, 2019.
  22. 22.Shih-Hung Liu, Shang-Yi Yu, Shao-Chi Wu, Hwann-Tzong Chen, and Tyng-Luh Liu. Learning gaussian instance segmentation in point clouds. arXiv:2007.09860, 2020.
  23. 23.Yongcheng Liu, Bin Fan, Shiming Xiang, and Chunhong Pan. Relation-shape convolutional neural network for point cloud analysis. In CVPR, 2019.
  24. 24.Ilya Loshchilov and Frank Hutter. Sgdr: Stochastic gradient descent with warm restarts. In ICVLR, 2017.
  25. 25.Daniel Maturana and Sebastian Scherer. Voxnet: A 3d convolutional neural network for real-time object recognition. In IROS, 2015.
  26. 26.Kaichun Mo, Shilin Zhu, Angel X Chang, Li Yi, Subarna Tripathi, Leonidas J Guibas, and Hao Su. Partnet: A large-scale benchmark for fine-grained and hierarchical part-level 3d object understanding. In CVPR, 2019.
  27. 27.Gaku Narita, Takashi Seno, Tomoya Ishikawa, and Yohsuke Kaji. Panopticfusion: Online volumetric semantic mapping at the level of stuff and things. arXiv:1903.01177, 2019.
  28. 28.Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer. Automatic differentiation in pytorch. In NIPS-W, 2017.
  29. 29.Quang-Hieu Pham, Thanh Nguyen, Binh-Son Hua, Gemma Roig, and Sai-Kit Yeung. Jsis3d: joint semantic-instance segmentation of 3d point clouds with multi-task pointwise networks and multi-value conditional random fields. In CVPR, 2019.
  30. 30.Charles R Qi, Or Litany, Kaiming He, and Leonidas J Guibas. Deep hough voting for 3d object detection in point clouds. In CVPR, 2019.
  31. 31.Charles R Qi, Wei Liu, Chenxia Wu, Hao Su, and Leonidas J Guibas. Frustum pointnets for 3d object detection from rgb-d data. In CVPR, 2018.
  32. 32.Charles R Qi, Hao Su, Kaichun Mo, and Leonidas J Guibas. Pointnet: Deep learning on point sets for 3d classification and segmentation. In CVPR, 2017.
  33. 33.Charles R Qi, Li Yi, Hao Su, and Leonidas J Guibas. Pointnet++: Deep hierarchical feature learning on point sets in a metric space. arXiv:1706.02413, 2017.
  34. 34.Gernot Riegler, Ali Osman Ulusoy, and Andreas Geiger. Octnet: Learning deep 3d representations at high resolutions. In CVPR, 2017.
  35. 35.Olaf Ronneberger, Philipp Fischer, and Thomas Brox. U-net: Convolutional networks for biomedical image segmentation. In MICCAI, 2015.
  36. 36.Radu Bogdan Rusu, Nico Blodow, and Michael Beetz. Fast point feature histograms (fpfh) for 3d registration. In ICVRA, 2009.
  37. 37.Radu Bogdan Rusu, Nico Blodow, Zoltan Csaba Marton, and Michael Beetz. Aligning point cloud views using persistent feature histograms. In IROS, 2008.
  38. 38.Yiru Shen, Chen Feng, Yaoqing Yang, and Dong Tian. Mining point cloud local structures by kernel correlation and graph pooling. In CVPR, 2018.
  39. 39.Martin Simonovsky and Nikos Komodakis. Dynamic edge-conditioned filters in convolutional neural networks on graphs. In CVPR, 2017.
  40. 40.Hugues Thomas, Charles R Qi, Jean-Emmanuel Deschaud, Beatriz Marcotegui, François Goulette, and Leonidas J Guibas. Kpconv: Flexible and deformable convolution for point clouds. In ICCV, 2019.
  41. 41.Weiyue Wang, Ronald Yu, Qiangui Huang, and Ulrich Neumann. Sgpn: Similarity group proposal network for 3d point cloud instance segmentation. In CVPR, 2018.
  42. 42.Xinlong Wang, Shu Liu, Xiaoyong Shen, Chunhua Shen, and Jiaya Jia. Associatively segmenting instances and semantics in point clouds. In CVPR, 2019.
  43. 43.Yue Wang, Yongbin Sun, Ziwei Liu, Sanjay E Sarma, Michael M Bronstein, and Justin M Solomon. Dynamic graph cnn for learning on point clouds. ACM TOG, 2019.
  44. 44.Wenxuan Wu, Zhongang Qi, and Li Fuxin. Pointconv: Deep convolutional networks on 3d point clouds. In CVPR, 2019.
  45. 45.Yifan Xu, Tianqi Fan, Mingye Xu, Long Zeng, and Yu Qiao. Spidercnn: Deep learning on point sets with parameterized convolutional filters. In ECCV, 2018.
  46. 46.Bo Yang, Jianan Wang, Ronald Clark, Qingyong Hu, Sen Wang, Andrew Markham, and Niki Trigoni. Learning object bounding boxes for 3d instance segmentation on point clouds. In NeurIPS, 2019.
  47. 47.Li Yi, Wang Zhao, He Wang, Minhyuk Sung, and Leonidas J Guibas. Gspn: Generative shape proposal network for 3d instance segmentation in point cloud. In CVPR, 2019.
  48. 48.Biao Zhang and Peter Wonka. Point cloud instance segmentation using probabilistic embeddings. In CVPR, 2021.
  49. 49.Biao Zhang and Peter Wonka. Point cloud instance segmentation using probabilistic embeddings. In CVPR, 2021.
  50. 50.Hengshuang Zhao, Li Jiang, Jiaya Jia, Philip HS Torr, and Vladlen Koltun. Point transformer. In ICCV, 2021.

Citation

MLA
Vu, T., et al. “SoftGroup for 3D Instance Segmentation on Point Clouds”. arXiv, 2022, http://arxiv.org/abs/2203.01509v1.
APA
Vu, T., Kim, K., Luu, T. M., Nguyen, X. T., & Yoo, C. D. (2022). SoftGroup for 3D Instance Segmentation on Point Clouds. arXiv. http://arxiv.org/abs/2203.01509v1
Chicago
Vu, T., K. Kim, T. M. Luu, X. T. Nguyen, and C. D. Yoo. 2022. “SoftGroup for 3D Instance Segmentation on Point Clouds”. arXiv. http://arxiv.org/abs/2203.01509v1.
Harvard
Vu, T. et al. (2022) “SoftGroup for 3D Instance Segmentation on Point Clouds”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2203.01509v1.
Vancouver
1. Vu T, Kim K, Luu TM, Nguyen XT, Yoo CD (2022) SoftGroup for 3D Instance Segmentation on Point Clouds. arXiv

BibTeX

@article{vu2022softgroup,
  title = {SoftGroup for 3D Instance Segmentation on Point Clouds},
  author = {Vu, Thang and Kim, Kookhoi and Luu, Tung M. and Nguyen, Xuan Thanh and Yoo, Chang D.},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2203.01509v1},
  eprint = {2203.01509}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE