ISBNet: a 3D Point Cloud Instance Segmentation Network with Instance-aware Sampling and Box-aware Dynamic Convolution

Tuan Duc NgoBinh-Son HuaKhoi Nguyen

article2023CVPR103 citations

Proposes a cluster-free 3D instance segmentation framework that couples instance-aware point sampling with box-guided dynamic convolutions to achieve state-of-the-art accuracy and fast inference on ScanNetV2, S3DIS, and STPLS3D benchmarks.

Listen

Three-dimensional instance segmentation—identifying individual objects and assigning semantic categories within 3D point cloud data—is a critical capability for technologies such as autonomous driving, augmented reality, and robotics. Existing bottom-up methods rely heavily on centroid clustering to group raw data points into distinct objects. However, these conventional approaches often fail when identical items are tightly packed together or when large, loosely connected objects are incorrectly fragmented into separate pieces.

The article introduces and evaluates ISBNet, a cluster-free framework designed to segment 3D point clouds accurately and efficiently. The primary objective is to demonstrate that combining instance-aware point sampling with box-aware dynamic convolution overcomes the clustering failures of prior systems and establishes superior performance across benchmark 3D datasets.

The authors evaluated ISBNet through empirical experiments on three standard indoor and outdoor datasets: ScanNetV2, S3DIS, and the aerial photogrammetry dataset STPLS3D. ISBNet replaces traditional grouping heuristics with an iterative sampling strategy that selects object candidate points while avoiding background and previously detected instances. These candidates aggregate local contextual features and jointly predict object bounding boxes. The network then integrates these predicted 3D bounding boxes as explicit geometric cues during dynamic convolution to isolate and generate final instance masks.

The findings show that ISBNet achieves state-of-the-art accuracy across all evaluated benchmarks while maintaining high processing speeds. First, ISBNet reached an average precision of 55.9 on the hidden ScanNetV2 test benchmark and exceeded the second-best method on the validation set by 3.7 points in average precision. Second, it surpassed existing state-of-the-art models on S3DIS cross-validation and STPLS3D, beating prior benchmarks by 3.4 and 3.0 points in average precision, respectively. Third, the proposed sampling technique achieved up to 100% instance candidate recall on ScanNetV2 validation data, compared to only 75.5% achieved by clustering methods. Finally, ISBNet demonstrated high computational efficiency, executing a complete scene analysis in 237 milliseconds on a single graphics processing unit, making it the fastest among evaluated state-of-the-art alternatives.

These results demonstrate that eliminating hand-tuned clustering in favor of direct instance-aware candidate sampling and bounding box geometric cues significantly enhances segmentation reliability. For real-world systems, this improvement reduces the risk of misidentifying overlapping obstacles or fragmenting complex assets, while the sub-second runtime supports the strict latency requirements necessary for safe, real-time autonomous navigation and mapping.

Organizations developing 3D perception pipelines should consider adopting cluster-free, dynamic convolution architectures over traditional clustering approaches when processing dense spatial data. Further development should focus on testing these models across additional real-world operational environments and exploring enhanced dynamic convolutions that incorporate richer geometric structures.

The confidence in these findings is strong across the evaluated standard datasets, though some limitations remain. The sampling step depends on the quality of intermediate instance predictions, meaning early errors could propagate through the pipeline. Additionally, axis-aligned bounding boxes do not always tightly fit irregularly shaped or intertwined objects, which can cause adjacent, touching objects—such as a counter and a refrigerator—to be erroneously merged.

arXiv: 2303.00246
Cover for ISBNet: a 3D Point Cloud Instance Segmentation Network with Instance-aware Sampling and Box-aware Dynamic Convolution

Abstract

Existing 3D instance segmentation methods are predomi-

nated by the bottom-up design – manually fine-tuned algo-

rithm to group points into clusters followed by a refinement

network. However, by relying on the quality of the clus-

ters, these methods generate susceptible results when (1)

nearby objects with the same semantic class are packed

together, or (2) large objects with loosely connected re-

gions. To address these limitations, we introduce ISBNet, a

novel cluster-free method that represents instances as ker-

nels and decodes instance masks via dynamic convolution.

To efficiently generate high-recall and discriminative ker-

nels, we propose a simple strategy named Instance-aware

Farthest Point Sampling to sample candidates and lever-

age the local aggregation layer inspired by PointNet++ to

encode candidate features. Moreover, we show that pre-

dicting and leveraging the 3D axis-aligned bounding boxes

in the dynamic convolution further boosts performance.

Our method set new state-of-the-art results on ScanNetV2

(55.9), S3DIS (60.8), and STPLS3D (49.2) in terms of AP

and retains fast inference time (237ms per scene on Scan-

NetV2). The source code and trained models are available at

https://github.com/VinAIResearch/ISBNet.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Our Approach
  • 3.1. Sampling-based Instance-wise Encoder
  • 3.2. Box-aware Dynamic Convolution
  • 3.3. Network Training
  • 4. Experiments
  • 4.1. Experimental Setup
  • 4.2. Main Results
  • 4.3. Qualitative Results
  • 4.4. Ablation Study
  • 5. Conclusion
  • References

Knowls

  1. Knowl 1 — ISBNet Architecture for 3D Instance Segmentation

    model/method

    ISBNet is a cluster-free 3D point cloud instance segmentation framework that represents object instances as dynamic convolution kernels and decodes instance masks using spatial and geometric cues.

    Given an input point cloud P∈RN×6\mathbf{P} \in \mathbb{R}^{N \times 6} with NN points (each represented by 3D coordinates (x,y,z)(x, y, z) and RGB color values), the architecture processes the scene through four core modules:

    1. 3D Backbone: A 3D U-Net built with sparse convolutions that extracts per-point feature representations F∈RN×D\mathbf{F} \in \mathbb{R}^{N \times D} (with feature dimension D=32D = 32).
    2. Point-wise Predictor: Multi-head prediction network taking per-point features F\mathbf{F} to generate point-wise semantic category probabilities, 3D axis-aligned bounding box predictions Fbox∈RN×6\mathbf{F}_{box} \in \mathbb{R}^{N \times 6} (parameterized by minimum and maximum coordinate bounds (x1,y1,z1,x2,y2,z2)(x_1, y_1, z_1, x_2, y_2, z_2)), and mask feature representations Fmask∈RN×H\mathbf{F}_{mask} \in \mathbb{R}^{N \times H} (H=32H = 32).
    3. Sampling-based Instance-wise Encoder: A sequence of Point Aggregator (PA) blocks that uses Instance-aware Farthest Point Sampling (IA-FPS) to select KK candidate points (K≪NK \ll N) and encodes their local context into instance-wise representations E∈RK×D\mathbf{E} \in \mathbb{R}^{K \times D}. Linear projection layers on E\mathbf{E} output instance classification scores L∈RK×C\mathbf{L} \in \mathbb{R}^{K \times C} (for CC semantic classes), instance-level bounding box predictions B∈RK×6\mathbf{B} \in \mathbb{R}^{K \times 6}, and instance convolution kernels W∈RK×H′\mathbf{W} \in \mathbb{R}^{K \times H'}.
    4. Box-aware Dynamic Convolution Head: Convolves the candidate-specific kernel w(k)\mathbf{w}^{(k)} with concatenated point features comprising point-wise mask features Fmask\mathbf{F}_{mask}, relative coordinate differences Fpos(k)\mathbf{F}_{pos}^{(k)}, and relative bounding box coordinate differences Fgeo(k)\mathbf{F}_{geo}^{(k)} to produce binary instance masks M^∈[0,1]K×N\widehat{\mathbf{M}} \in [0, 1]^{K \times N}.
  2. Knowl 2 — Box-aware Dynamic Convolution Formulation

    equation

    Box-aware Dynamic Convolution predicts an instance binary mask m^(k)∈[0,1]1×N\widehat{\mathbf{m}}^{(k)} \in [0, 1]^{1 \times N} for the kk-th instance query candidate across all NN points by conditioning the convolution on both appearance and 3D bounding box geometry:

    ight]; \mathbf{w}^{(k)}\right)\right)$$ where: - $[\cdot ; \cdot]$ denotes channel-wise concatenation. - $\mathbf{F}_{mask} \in \mathbb{R}^{N \times H}$ is the point-wise mask feature map output by the point-wise predictor ($H = 32$). - $\mathbf{F}_{pos}^{(k)} \in \mathbb{R}^{N \times 3}$ represents the relative 3D coordinate displacement of every point $\mathbf{p}^{(i)} \in \mathbb{R}^3$ relative to candidate point $k$: $\mathbf{f}_{pos}^{(k, i)} = \mathbf{p}^{(i)} - \mathbf{p}^{(k)}$. - $\mathbf{F}_{geo}^{(k)} \in \mathbb{R}^{N \times 6}$ is the geometric box feature defined by the element-wise absolute difference between the predicted 6D axis-aligned bounding box $\mathbf{f}_{box}^{(i)} = (x_1, y_1, z_1, x_2, y_2, z_2)$ of point $i$ and the candidate's predicted bounding box $\mathbf{f}_{box}^{(k)}$: $$\mathbf{f}_{geo}^{(k, i)} = \left| \mathbf{f}_{box}^{(i)} - \mathbf{f}_{box}^{(k)} \right|$$ - $\mathbf{w}^{(k)} \in \mathbb{R}^{H'}$ is the dynamic convolution kernel generated for candidate $k$. - $\text{Conv}(\cdot; \mathbf{w}^{(k)})$ is implemented as a multi-layer dynamic convolutional network (with 2 layers and hidden channel dimension 32) parameterized by $\mathbf{w}^{(k)}$.
  3. Knowl 3 — Instance-aware Farthest Point Sampling

    model/method

    Instance-aware Farthest Point Sampling (IA-FPS) is a candidate selection strategy designed to maximize instance recall across varying object scales by filtering out background points and points belonging to previously identified instances.

    Given an input point set P\mathbf{P}, IA-FPS defines the candidate sampling domain P′⊂P\mathbf{P}' \subset \mathbf{P} as:

    P′={p(i)∈P  |  min⁡k=0…K′(1−m(k)(i))>τ}\mathbf{P}' = \left\{ \mathbf{p}^{(i)} \in \mathbf{P} \;\middle|\; \min_{k=0 \dots K'} \left(1 - m_{(k)}^{(i)}\right) > \tau \right\}

    where m(0)(i)∈[0,1]m_{(0)}^{(i)} \in [0, 1] is the estimated background probability for point ii from point-wise semantic prediction, m(k)(i)∈[0,1]m_{(k)}^{(i)} \in [0, 1] is the predicted instance mask probability of point ii generated by the kk-th candidate, K′K' is the number of already selected candidate queries, and τ\tau is a foreground threshold (set to τ=0.5\tau = 0.5). Standard Farthest Point Sampling is then executed over P′\mathbf{P}'.

    In training, because early mask predictions are unrefined, all K=256K = 256 candidates are sampled concurrently from the predicted foreground mask satisfying 1−m(0)(i)>τ1 - m_{(0)}^{(i)} > \tau.

    In inference, candidates are sampled iteratively in sequential chunks {κ1,…,κT}\{\kappa_1, \dots, \kappa_T\} (e.g., chunks of sizes 192, 128, and 64 for a total K=384K = 384). Subsequent chunks are sampled exclusively from points that are neither background nor covered by the predicted masks of earlier chunks. On ScanNetV2 validation, IA-FPS achieves instance recall of 100%, 98.4%, 94.5%, and 89.2% for 2048, 512, 256, and 128 sample points, respectively, compared to standard FPS (99.3%, 93.3%, 85.4%, 71.3%) and centroid clustering (75.5%).

  4. Knowl 4 — Point Aggregator Block for Local Context Encoding

    model/method

    The Point Aggregator (PA) block encodes local geometric and semantic context around sampled candidate points into instance-level representations without relying on centroid-based point clustering.

    For each sampled candidate kk, a ball-query collects Q=32Q = 32 spatial neighbors within search radius rr. The local feature representation e(k)∈RD\mathbf{e}^{(k)} \in \mathbb{R}^D is computed via:

    e(k)=f(k)+max⁡q(MLP([Flocal(k);Plocal(k)]))\mathbf{e}^{(k)} = \mathbf{f}^{(k)} + \max_q \left( \text{MLP}\left( \left[ \mathbf{F}_{local}^{(k)} ; \mathbf{P}_{local}^{(k)} \right] \right) \right)

    where:

    • f(k)∈RD\mathbf{f}^{(k)} \in \mathbb{R}^D is the candidate's original backbone feature, incorporated via a residual connection to prevent vanishing gradients.
    • Flocal(k)∈RQ×D\mathbf{F}_{local}^{(k)} \in \mathbb{R}^{Q \times D} is the matrix of gathered local neighbor features.
    • Plocal(k)∈[−1,1]Q×3\mathbf{P}_{local}^{(k)} \in [-1, 1]^{Q \times 3} represents the relative neighbor coordinates normalized by the search radius rr: (pq−pk)/r(\mathbf{p}_q - \mathbf{p}_k) / r.
    • [⋅;⋅][\cdot ; \cdot] denotes concatenation, and max⁡q\max_q performs channel-wise max-pooling over the QQ neighbors.

    To expand the receptive field hierarchically, the instance-wise encoder stacks two successive PA blocks. The first block operates with ball-query radius r=0.2 mr = 0.2\text{ m}, and the second block samples from the subset generated by the first with radius r=0.4 mr = 0.4\text{ m}.

  5. Knowl 5 — Multi-Task Loss and One-to-Many Hungarian Matching

    equation

    The training objective of ISBNet optimizes both point-wise predictions and instance-level predictions. The total loss consists of a point-wise loss and an instance-wise loss.

    The point-wise loss includes cross-entropy loss for semantic segmentation, and L1L_1 plus generalized Intersection-over-Union (gIoU) loss for 3D axis-aligned bounding box regression.

    For instance-wise loss, candidate predictions are matched to ground-truth instances using a one-to-many Hungarian matching scheme where each ground-truth instance is duplicated S=4S = 4 times. The matching cost between predicted candidate kk and ground-truth instance jj is:

    C(k,j)=γmaskCmask(m^(k),m(j))+Ccls(l^(k),l(j))C(k, j) = \gamma_{mask} C_{mask}\left(\widehat{\mathbf{m}}^{(k)}, \mathbf{m}^{(j)}\right) + C_{cls}\left(\widehat{\mathbf{l}}^{(k)}, \mathbf{l}^{(j)}\right)

    where CmaskC_{mask} is the Dice loss between the predicted binary mask m^(k)\widehat{\mathbf{m}}^{(k)} and ground-truth binary mask m(j)\mathbf{m}^{(j)}, CclsC_{cls} is classification cross-entropy, and γmask=5\gamma_{mask} = 5.

    Following assignment, the instance-wise loss is computed as:

    Linst=Lcls+λboxLbox+λmaskLmask+λmsLMSL_{inst} = L_{cls} + \lambda_{box} L_{box} + \lambda_{mask} L_{mask} + \lambda_{ms} L_{MS}

    where:

    • LclsL_{cls} is instance classification cross-entropy loss.
    • LboxL_{box} combines L1L_1 loss and gIoU loss on axis-aligned bounding box coordinates, weighted by λbox=1\lambda_{box} = 1.
    • LmaskL_{mask} combines Dice loss and Binary Cross-Entropy (BCE) loss on predicted masks, weighted by λmask=5\lambda_{mask} = 5.
    • LMSL_{MS} is Mask-Scoring loss, weighted by λms=1\lambda_{ms} = 1.
  6. Knowl 6 — 3D Point Cloud Instance Segmentation Benchmark Performance

    data/table

    ISBNet was evaluated on three standard 3D instance segmentation benchmarks: ScanNetV2 (indoor mesh scans, evaluated on test and validation sets), S3DIS (indoor scenes, evaluated under Area 5 and 6-fold cross-validation), and STPLS3D (synthetic and real aerial photogrammetry point clouds). Metrics include Average Precision (AP), AP at 50% IoU (AP50), AP at 25% IoU (AP25), 3D Box AP50, mean coverage (mCov), mean weighted coverage (mWCov), mean precision (mPrec50), and mean recall (mRec50).

    Benchmark AP AP50 AP25 Box AP50 mCov mRec50
    ScanNetV2 Test 55.9 76.3 84.5 - - -
    ScanNetV2 Val 54.5 73.1 82.5 62.0 - -
    S3DIS (Area 5) 56.3 67.5 - - 70.0 72.0
    S3DIS (6-fold) 60.8 70.5 - - 74.9 77.1
    STPLS3D Val 49.2 64.0 - - - -

    On the ScanNetV2 test benchmark, ISBNet achieves 55.9% AP, outperforming previous methods including DKNet (53.2%), SSTNet (50.6%), and SoftGroup (50.4%). On ScanNetV2 validation, ISBNet reaches 54.5% AP and 62.0% Box AP50, surpassing SoftGroup (46.0% AP / 59.4% Box AP50) and DKNet (50.8% AP / 59.0% Box AP50). On S3DIS 6-fold cross-validation, ISBNet achieves 60.8% AP and 74.9% mCov, outperforming SoftGroup (54.4% AP / 69.3% mCov).

  7. Knowl 7 — Inference Latency and Component Runtime Breakdown

    empirical result

    Inference runtime was benchmarked on ScanNetV2 validation scenes using a single NVIDIA Titan X GPU. Processing is decomposed into three execution stages: 3D backbone feature extraction, instance abstraction (query candidate generation and encoding), and mask decoding (dynamic convolution).

    ISBNet achieves a total per-scene inference latency of 237 ms, decomposed as:

    • Backbone: 152 ms
    • Instance Abstractor (IA-FPS + Point Aggregator): 53 ms
    • Mask Decoder (Box-aware Dynamic Convolution): 32 ms

    In comparison, total runtimes and validation AP of prior methods on the identical hardware are:

    • PointGroup: 452 ms (34.8 AP)
    • HAIS: 339 ms (43.5 AP)
    • SoftGroup: 345 ms (46.0 AP)
    • DyCo3D: 302 ms (40.6 AP)
    • DKNet: 614 ms (50.8 AP)

    Replacing clustering-based instance grouping (which requires centroid-offset clustering and voxelization) with Point Aggregator blocks reduces the instance abstraction latency to 53 ms while achieving 54.5 AP.

  8. Knowl 8 — Ablation on ISBNet Architectural Modules and Bounding Box Supervision

    data/table

    Ablation experiments on the ScanNetV2 validation set isolate the impact of Instance-aware Farthest Point Sampling (IA-FPS), the Local Aggregation Layer (LAL), and Box-aware Dynamic Convolution (BA-DyCo), as well as the dual role of 3D axis-aligned bounding box prediction.

    Configuration IA-FPS LAL BA-DyCo AP AP50 AP25
    DyCo3D* (Re-impl.) - - - 49.4 67.6 77.4
    Baseline (FPS only) - - - 47.9 66.4 77.1
    + IA-FPS ✓ - - 49.7 67.5 78.6
    + LAL - ✓ - 50.1 69.4 79.1
    + BA-DyCo - - ✓ 48.6 67.7 77.8
    + IA-FPS + LAL ✓ ✓ - 53.4 71.9 81.8
    ISBNet (Full) ✓ ✓ ✓ 54.5 73.1 82.5

    Evaluating the bounding box prediction mechanisms reveals distinct contributions from multi-task loss and geometric conditioning:

    • No Box Supervision: 52.8 AP / 71.6 AP50
    • Auxiliary Box Loss only (no geometric feature in dynamic conv): 53.4 AP / 71.9 AP50 (+0.6 AP)
    • Auxiliary Box Loss + Geometric Box Feature (Fgeo(k)F_{geo}^{(k)} in dynamic conv): 54.5 AP / 73.1 AP50 (+1.1 AP over loss-only, +1.7 AP overall)
  9. Knowl 9 — Ablation on IA-FPS Iterative Chunking and Dynamic Convolution Design

    empirical result

    Ablation on the ScanNetV2 validation set examines the effects of candidate chunking schedules in IA-FPS and structural choices in dynamic convolution:

    1. IA-FPS Chunk Schedules (KK total candidates):

      • Single chunk K=256K = 256: 53.9 AP / 72.2 AP50
      • Single chunk K=384K = 384: 54.2 AP / 72.4 AP50
      • Single chunk K=512K = 512: 53.6 AP / 71.9 AP50
      • Iterative equal chunks (128,128,128)(128, 128, 128) (K=384K = 384): 54.0 AP / 72.8 AP50
      • Iterative decaying chunks (192,128,64)(192, 128, 64) (K=384K = 384): 54.5 AP / 73.1 AP50 Iterative sampling with decreasing chunk sizes prevents point redundancy and ensures superior coverage.
    2. Dynamic Convolution Architecture:

      • 1 layer, dimensions (41,1)(41, 1), 41 parameters: 45.7 AP / 67.1 AP50
      • 2 layers, dimensions (25,8,1)(25, 8, 1), 216 parameters: 53.6 AP / 72.1 AP50
      • 2 layers, dimensions (41,16,1)(41, 16, 1), 688 parameters: 53.9 AP / 72.3 AP50
      • 2 layers, dimensions (41,32,1)(41, 32, 1), 1376 parameters: 54.5 AP / 73.1 AP50
      • 3 layers, dimensions (41,16,16,1)(41, 16, 16, 1), 960 parameters: 53.9 AP / 72.7 AP50
      • 3 layers, dimensions (41,32,16,1)(41, 32, 16, 1), 1696 parameters: 54.2 AP / 72.8 AP50 A 2-layer dynamic convolution head with 32 hidden channels achieves the optimal trade-off between capacity and generalization.
    3. Number of Stacked Point Aggregator (PA) Blocks:

      • 1 PA block: 53.2 AP / 72.5 AP50
      • 2 PA blocks: 54.5 AP / 73.1 AP50
      • 3 PA blocks: 54.3 AP / 73.0 AP50
  10. Knowl 10 — Limitations of Axis-Aligned Bounding Box Condition and Iterative Sampling

    limitation

    ISBNet exhibits two primary limitations:

    1. Coincident or Enclosed Bounding Boxes: Axis-aligned 3D bounding boxes cannot tightly wrap irregular or nested geometric structures. When distinct object instances share overlapping or nearly identical bounding boxes (such as a refrigerator embedded directly inside a surrounding kitchen counter), the geometric difference feature Fgeo(k)\mathbf{F}_{geo}^{(k)} becomes non-discriminative, causing the model to merge points from different instances into a single object mask.
    2. Sequential Sampling Error Propagation in IA-FPS: The iterative candidate selection mechanism relies on the accuracy of instance masks predicted in earlier chunks to mask out points for subsequent sampling rounds. If early mask predictions miss foreground regions or erroneously classify foreground points as background, those regions are omitted from candidate sampling, preventing recovery in subsequent chunks.

Coverage note — None was omitted. All major contributed architectural concepts, mathematical formulations, training loss designs, empirical results across benchmarks, runtime analyses, ablation studies, and stated limitations are fully covered.

References

  1. 1.I. Armeni, S. Sax, A. R. Zamir, and S. Savarese. Joint 2d-3d-semantic data for indoor scene understanding. arXiv preprint arXiv:1702.01105, 2017.
  2. 2.Z. Cai and N. Vasconcelos. Cascade R-CNN: Delving into high quality object detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018.
  3. 3.N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kirillov, and S. Zagoruyko. End-to-end object detection with transformers. In Proceedings of the European Conference on Computer Vision, 2020.
  4. 4.M. Chen, Q. Hu, Z. Yu, H. THOMAS, A. Feng, Y. Hou, K. McCullough, F. Ren, and L. Soibelman. Stpls3d: A large-scale synthetic and real aerial photogrammetry 3d point cloud dataset. In Proceedings of the British Machine Vision Conference, 2022.
  5. 5.S. Chen, J. Fang, Q. Zhang, W. Liu, and X. Wang. Hierarchical aggregation for 3d instance segmentation. In Proceedings of the International Conference on Computer Vision, 2021.
  6. 6.B. Cheng, A. G. Schwing, and A. Kirillov. Per-pixel classification is not all you need for semantic segmentation. In Advances in Neural Information Processing Systems, 2021.
  7. 7.B. Cheng, I. Misra, A. G. Schwing, A. Kirillov, and R. Girdhar. Masked-attention mask transformer for universal image segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2022.
  8. 8.A. Dai, A. X. Chang, M. Savva, M. Halber, T. Funkhouser, and M. Nießner. Scannet: Richly-annotated 3d reconstructions of indoor scenes. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2017.
  9. 9.S. Dong, G. Lin, and T.-Y. Hung. Learning regional purity for instance segmentation on 3d point clouds. In Proceedings of the European Conference on Computer Vision, 2022.
  10. 10.Y. Eldar, M. Lindenbaum, M. Porat, and Y. Y. Zeevi. The farthest point strategy for progressive image sampling. IEEE Transactions on Image Processing, 6(9):1305–1315, 1997.
  11. 11.F. Engelmann, M. Bokeloh, A. Fathi, B. Leibe, and M. Nießner. 3d-mpa: Multi-proposal aggregation for 3d semantic instance segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2020.
  12. 12.B. Graham, M. Engelcke, and L. Van Der Maaten. 3d semantic segmentation with submanifold sparse convolutional networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018.
  13. 13.R. Guo, D. Niu, L. Qu, and Z. Li. Sotr: Segmenting objects with transformers. In Proceedings of the International Conference on Computer Vision, 2021.
  14. 14.L. Han, T. Zheng, L. Xu, and L. Fang. Occuseg: Occupancy-aware 3d instance segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2020.
  15. 15.K. He, G. Gkioxari, P. Dollár, and R. Girshick. Mask r-cnn. In Proceedings of the International Conference on Computer Vision, 2017.
  16. 16.T. He, C. Shen, and A. van den Hengel. Dyco3d: Robust instance segmentation of 3d point clouds through dynamic convolution. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2021.
  17. 17.T. He, C. Shen, and A. van den Hengel. Pointinst3d: Segmenting 3d instances by points. In Proceedings of the European Conference on Computer Vision, 2022.
  18. 18.J. Hou, A. Dai, and M. Nießner. 3d-sis: 3d semantic instance segmentation of rgb-d scans. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2019.
  19. 19.Z. Huang, L. Huang, Y. Gong, C. Huang, and X. Wang. Mask scoring r-cnn. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2019.
  20. 20.D. Jia, Y. Yuan, H. He, X. Wu, H. Yu, W. Lin, L. Sun, C. Zhang, and H. Hu. Detrs with hybrid matching. arXiv preprint arXiv:2207.13080, 2022.
  21. 21.L. Jiang, H. Zhao, S. Shi, S. Liu, C.-W. Fu, and J. Jia. Pointgroup: Dual-set point grouping for 3d instance segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2020.
  22. 22.A. Kirillov, Y. Wu, K. He, and R. Girshick. Pointrend: Image segmentation as rendering. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2020.
  23. 23.L. Landrieu and M. Boussaha. Point cloud oversegmentation with graph-structured deep metric learning. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2019.
  24. 24.L. Landrieu and M. Simonovsky. Large-scale point cloud semantic segmentation with superpoint graphs. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018.
  25. 25.F. Li, H. Zhang, S. Liu, L. Zhang, L. M. Ni, H.-Y. Shum, et al. Mask dino: Towards a unified transformer-based framework for object detection and segmentation. arXiv preprint arXiv:2206.02777, 2022.
  26. 26.Z. Liang, Z. Li, S. Xu, M. Tan, and K. Jia. Instance segmentation in 3d scenes using semantic superpoint tree networks. In Proceedings of the International Conference on Computer Vision, 2021.
  27. 27.A. Paszke, S. Gross, S. Chintala, G. Chanan, E. Yang, Z. DeVito, Z. Lin, A. Desmaison, L. Antiga, and A. Lerer. Automatic differentiation in pytorch. 2017.
  28. 28.Q.-H. Pham, T. Nguyen, B.-S. Hua, G. Roig, and S.-K. Yeung. Jsis3d: Joint semantic-instance segmentation of 3d point clouds with multi-task pointwise networks and multi-value conditional random fields. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2019.
  29. 29.C. R. Qi, L. Yi, H. Su, and L. J. Guibas. Pointnet++: Deep hierarchical feature learning on point sets in a metric space. In Advances in Neural Information Processing Systems, 2017.
  30. 30.S. Ren, K. He, R. Girshick, and J. Sun. Faster r-cnn: Towards real-time object detection with region proposal networks. In Advances in Neural Information Processing Systems, 2015.
  31. 31.H. Rezatofighi, N. Tsoi, J. Gwak, A. Sadeghian, I. Reid, and S. Savarese. Generalized intersection over union. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2019.
  32. 32.C. H. Sudre, W. Li, T. Vercauteren, S. Ourselin, and M. J. Cardoso. Generalised dice overlap as a deep learning loss function for highly unbalanced segmentations. In Deep learning in medical image analysis and multimodal learning for clinical decision support. Springer, 2017.
  33. 33.Z. Tian, C. Shen, and H. Chen. Conditional convolutions for instance segmentation. In Proceedings of the European Conference on Computer Vision. Springer, 2020.
  34. 34.T. Vu, K. Kim, T. M. Luu, X. T. Nguyen, and C. D. Yoo. Softgroup for 3d instance segmentation on 3d point clouds. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2022.
  35. 35.W. Wang, R. Yu, Q. Huang, and U. Neumann. Sgpn: Similarity group proposal network for 3d point cloud instance segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018.
  36. 36.X. Wang, T. Kong, C. Shen, Y. Jiang, and L. Li. Solo: Segmenting objects by locations. In Proceedings of the European Conference on Computer Vision. Springer, 2020.
  37. 37.X. Wang, R. Zhang, T. Kong, L. Li, and C. Shen. Solov2: Dynamic and fast instance segmentation. In Advances in Neural Information Processing Systems, 2020.
  38. 38.Y. Wu, M. Shi, S. Du, H. Lu, Z. Cao, and W. Zhong. 3d instances as 1d kernels. In Proceedings of the European Conference on Computer Vision, 2022.
  39. 39.B. Yang, J. Wang, R. Clark, Q. Hu, S. Wang, A. Markham, and N. Trigoni. Learning object bounding boxes for 3d instance segmentation on point clouds. In Advances in Neural Information Processing Systems, 2019.
  40. 40.L. Yi, W. Zhao, H. Wang, M. Sung, and L. J. Guibas. Gspn: Generative shape proposal network for 3d instance segmentation in point cloud. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2019.
  41. 41.B. Zhang and P. Wonka. Point cloud instance segmentation using probabilistic embeddings. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2021.
  42. 42.W. Zhao, Y. Yan, C. Yang, J. Ye, X. Yang, and K. Huang. Divide and conquer: 3d point cloud instance segmentation with point-wise binarization. In Proceedings of the European Conference on Computer Vision, 2022.

Citation

MLA
Ngo, T. D., et al. “ISBNet: A 3D Point Cloud Instance Segmentation Network with Instance-aware Sampling and Box-aware Dynamic Convolution”. arXiv, 2023, http://arxiv.org/abs/2303.00246v2.
APA
Ngo, T. D., Hua, B.-S., & Nguyen, K. (2023). ISBNet: a 3D Point Cloud Instance Segmentation Network with Instance-aware Sampling and Box-aware Dynamic Convolution. arXiv. http://arxiv.org/abs/2303.00246v2
Chicago
Ngo, T. D., B.-S. Hua, and K. Nguyen. 2023. “ISBNet: A 3D Point Cloud Instance Segmentation Network with Instance-aware Sampling and Box-aware Dynamic Convolution”. arXiv. http://arxiv.org/abs/2303.00246v2.
Harvard
Ngo, T.D., Hua, B.-S. and Nguyen, K. (2023) “ISBNet: a 3D Point Cloud Instance Segmentation Network with Instance-aware Sampling and Box-aware Dynamic Convolution”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2303.00246v2.
Vancouver
1. Ngo TD, Hua B-S, Nguyen K (2023) ISBNet: a 3D Point Cloud Instance Segmentation Network with Instance-aware Sampling and Box-aware Dynamic Convolution. arXiv

BibTeX

@article{ngo2023isbnet,
  title = {ISBNet: a 3D Point Cloud Instance Segmentation Network with Instance-aware Sampling and Box-aware Dynamic Convolution},
  author = {Ngo, Tuan Duc and Hua, Binh-Son and Nguyen, Khoi},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2303.00246v2},
  eprint = {2303.00246}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE