CoDA: Collaborative Novel Box Discovery and Cross-modal Alignment for Open-vocabulary 3D Object Detection

Yang CaoYihan ZengHang XuDan Xu

article2023NeurIPS69 citations

Proposes an end-to-end framework that couples 3D novel box discovery with cross-modal alignment to simultaneously localize and classify unseen 3D objects without relying on external 2D open-vocabulary detectors.

Listen

Three-dimensional (3D) object detection plays a vital role in real-world technologies such as autonomous driving, robotics, and industrial manufacturing. However, traditional systems are restricted to fixed, pre-defined sets of categories because annotating 3D point cloud data is extremely labor-intensive and costly. Open-vocabulary 3D object detection addresses this limitation by enabling models to recognize and localize novel, unbounded categories. The core challenge lies in simultaneously identifying the location of unseen objects and classifying them accurately when only a very small set of base categories has human annotations.

The main objective of the article is to demonstrate a unified framework, named CoDA, that simultaneously discovers novel 3D bounding boxes and classifies unseen objects without relying on complex, external two-dimensional (2D) open-vocabulary detection models. It evaluates the performance of this framework against existing baseline and alternative detection methods across benchmark indoor datasets.

To achieve this, the article introduces a collaborative approach combining a 3D Novel Object Discovery strategy with a Discovery-Driven Cross-Modal Alignment module. The novel object discovery strategy leverages 3D geometric bounding box patterns from known base categories alongside 2D semantic priors from the pre-trained CLIP vision-language model to propose and verify pseudo bounding boxes for unseen objects. The cross-modal alignment module then aligns 3D point cloud features with 2D image and text representations through category-agnostic distillation and class-specific contrastive learning. The discovery of novel boxes and the alignment of multi-modal features are trained iteratively, where improvements in one directly enhance the other. The framework was evaluated on two widely used benchmark datasets: SUN-RGBD (using 10 base and 36 novel categories) and ScanNet (using 10 base and 50 novel categories).

The findings show substantial improvements over previous techniques. First, the unified CoDA framework outperformed the best-performing alternative methods by more than 80% in mean Average Precision (mAP) for 3D object detection. Second, the novel object discovery mechanism increased novel category recall from approximately 21.5% in the baseline to 33.7% in the complete framework on the SUN-RGBD dataset, overcoming category forgetting during extended training. Third, collaborative learning between object discovery and cross-modal feature alignment improved precision on novel categories by roughly 26% and base categories by 48% compared to discovery alone. Fourth, during deployment and testing, the framework successfully detected and categorized novel objects using purely 3D point cloud data without requiring 2D camera images.

These results demonstrate that systems can successfully adapt to open-world environments with minimal manual annotation, significantly reducing the cost and timeline associated with data labeling. Unlike prior approaches that require heavy secondary 2D detectors, the proposed method provides an end-to-end pathway that reduces system dependencies while maintaining high detection accuracy. For robotics and autonomous systems, this expands operational safety and flexibility in novel, unstructured environments.

Organizations developing spatial computing, robotics, or autonomous vehicle systems should consider incorporating collaborative 3D-to-2D feature alignment architectures to handle unknown object classes. Technical teams should run pilot evaluations on their domain-specific point cloud data to determine if discovery-driven distillation can reduce manual labeling overhead. Because current evaluations remain focused on indoor benchmark environments, further validation on outdoor datasets and under noisy sensor conditions is recommended before full-scale commercial deployment.

  • Paper: WildDet3D: Scaling Promptable 3D Detection in the Wild, Weikai Huang et al. (2026). WildDet3D scales promptable, open-vocabulary 3D detection to in-the-wild settings across thousands of categories, generalizing beyond the indoor benchmark discoveries developed in CoDA.
Cover for CoDA: Collaborative Novel Box Discovery and Cross-modal Alignment for Open-vocabulary 3D Object Detection

Abstract

Open-vocabulary 3D Object Detection (OV-3DDet) aims to detect objects from an arbitrary list of categories within a 3D scene, which remains seldom explored in the literature. There are primarily two fundamental problems in OV-3DDet, i.e., localizing and classifying novel objects. This paper aims at addressing the two problems simultaneously via a unified framework, under the condition of limited base categories. To localize novel 3D objects, we propose an effective 3D Novel Object Discovery strategy, which utilizes both the 3D box geometry priors and 2D semantic open-vocabulary priors to generate pseudo box labels of the novel objects. To classify novel object boxes, we further develop a cross-modal alignment module based on discovered novel boxes, to align feature spaces between 3D point cloud and image/text modalities. Specifically, the alignment process contains a class-agnostic and a class-discriminative alignment, incorporating not only the base objects with annotations but also the increasingly discovered novel objects, resulting in an iteratively enhanced alignment. The novel box discovery and cross-modal alignment are jointly learned to collaboratively benefit each other. The novel object discovery can directly impact the cross-modal alignment, while a better feature alignment can, in turn, boost the localization capability, leading to a unified OV-3DDet framework, named CoDA, for simultaneous novel object localization and classification. Extensive experiments on two challenging datasets (i.e., SUN-RGBD and ScanNet) demonstrate the effectiveness of our method and also show a significant mAP improvement upon the best-performing alternative method by 80%. Codes and pre-trained models are released on the project page¹.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Methods
  • 3.1 Framework Overview
  • 3.2 3D Novel Object Discovery (3D-NOD) with Cross-Priors
  • 3.3 Discovery-Driven Cross-Modal Alignment (DCMA)
  • 4 Experiments
  • 4.1 Experimental Setup
  • 4.2 Model Analysis
  • 4.3 Comparison with Alternatives
  • 5 Conclusion
  • References

Knowls

  1. Knowl 1 — CoDA unified open-vocabulary 3D detector

    model/method

    CoDA is an end-to-end open-vocabulary 3D object detector that learns to localize and classify categories not annotated during training. Its backbone is a transformer-based 3DETR encoder–decoder with object queries, a 3D box localization head, and an object classification head. CoDA adds two jointly optimized components: 3D Novel Object Discovery (3D-NOD), which generates pseudo-labels for novel 3D boxes, and Discovery-Driven Cross-Modal Alignment (DCMA), which aligns point-cloud object features with CLIP image and text features. The model uses human annotations only for a limited set of base categories and does not require an external open-vocabulary 2D detector. CLIP images and text provide training-time supervision, while CoDA requires only a point cloud as input at test time.

  2. Knowl 2 — Cross-prior 3D Novel Object Discovery

    model/method

    3D-NOD discovers novel 3D boxes by combining class-agnostic 3D geometry with open-vocabulary 2D semantic evidence. Let CSeenC^{\mathrm{Seen}} be the annotated base-category set, let ℓj\ell_j be a labeled 3D box, and let cjc_j be its category. The initial annotated box pool is

    O0base={oj=(ℓj,cj)∣cj∈CSeen}.O^{\mathrm{base}}_0=\{o_j=(\ell_j,c_j)\mid c_j\in C^{\mathrm{Seen}}\}.

    A detector W0W_0 is first trained on this pool using 3D box regression and binary objectness, but without a base-category classification loss. For object query nn, the detector predicts a 3D box ℓn3D\ell^{3D}_n and objectness probability pngp^g_n. The box is projected into the corresponding color image using the camera intrinsic matrix MM:

    ℓn2D=M×ℓn3D.\ell^{2D}_n=M\times \ell^{3D}_n.

    The projected image region is cropped and encoded by the CLIP image encoder as FI,nObjF^{\mathrm{Obj}}_{I,n}. A predefined super-category vocabulary with CC text categories is encoded by the CLIP text encoder as FTSuperF^{\mathrm{Super}}_T. The semantic distribution over those categories is

    Pn3DObj=Softmax⁡(FI,nObj⋅FTSuper)={pn,1s,…,pn,Cs},P^{3D\mathrm{Obj}}_n=\operatorname{Softmax}\left(F^{\mathrm{Obj}}_{I,n}\cdot F^{\mathrm{Super}}_T\right)=\{p^s_{n,1},\ldots,p^s_{n,C}\},

    where ⋅\cdot is the dot product and pn,isp^s_{n,i} is the semantic probability for category ii. Let cn∗=arg⁡max⁡cpn,csc_n^*=\arg\max_c p^s_{n,c}, and let θg\theta_g and θs\theta_s be geometry and semantic confidence thresholds. A candidate is accepted as a novel object when it has low 3D overlap with every annotated base box, sufficient 3D objectness, sufficient CLIP semantic confidence, and a predicted category outside CSeenC^{\mathrm{Seen}}:

    Otdisc={on | IoU⁡3D(on,oi′)<0.25 ∀oi′∈O0base, png>θg, pn,cn∗s>θs, cn∗∉CSeen}.O^{\mathrm{disc}}_t=\left\{o_n\ \middle|\ \operatorname{IoU}_{3D}(o_n,o_i')<0.25\ \forall o_i'\in O^{\mathrm{base}}_0,\ p^g_n>\theta_g,\ p^s_{n,c_n^*}>\theta_s,\ c_n^*\notin C^{\mathrm{Seen}}\right\}.

    The discovered boxes are accumulated rather than replaced:

    Ot+1novel=Otnovel∪Otdisc.O^{\mathrm{novel}}_{t+1}=O^{\mathrm{novel}}_t\cup O^{\mathrm{disc}}_t.

    The expanded pool supplies supervision to later detector updates, allowing the current 3D detector and the CLIP semantic prior to discover progressively more novel objects.

  3. Knowl 3 — Class-agnostic 3D-to-2D feature distillation

    equation

    DCMA aligns each predicted 3D object-query feature with the CLIP feature of its projected image region without using a category label. For query n∈{1,…,N}n\in\{1,\ldots,N\}, where NN is the number of object queries, let Fn3DObjF^{3D\mathrm{Obj}}_n be the final decoder feature and let Fn2DObjF^{2D\mathrm{Obj}}_n be the CLIP image feature obtained by projecting the predicted 3D box into the color image and cropping the corresponding 2D region. The class-agnostic distillation loss is

    Ldistill3D↔2D=∑n=1N∥Fn3DObj−Fn2DObj∥1.L^{3D\leftrightarrow 2D}_{\mathrm{distill}}=\sum_{n=1}^{N}\left\|F^{3D\mathrm{Obj}}_n-F^{2D\mathrm{Obj}}_n\right\|_1.

    This loss transfers open-world visual information from CLIP to point-cloud features wherever the detector predicts a 3D box. It does not require category annotations and is therefore applicable to both annotated base objects and unlabeled or discovered objects; its objective is modality alignment rather than category discrimination.

  4. Knowl 4 — Discovery-driven text contrastive alignment

    equation

    For each 3D object-query feature Fn3DObjF^{3D\mathrm{Obj}}_n, DCMA computes similarities to the CLIP text embeddings FTSuperF^{\mathrm{Super}}_T of the super-category vocabulary:

    Sn=Softmax⁡(Fn3DObj⋅FTSuper).S_n=\operatorname{Softmax}\left(F^{3D\mathrm{Obj}}_n\cdot F^{\mathrm{Super}}_T\right).

    The predicted 3D boxes are bipartite-matched to boxes in a label pool OlabelO^{\mathrm{label}} containing foreground base or discovered novel boxes. For a matched query, the associated category produces a one-hot target vector hnh_n; an unmatched query is excluded so that background or noisy boxes do not receive category supervision. With 1(Fn3DObj,Olabel)\mathbf{1}(F^{3D\mathrm{Obj}}_n,O^{\mathrm{label}}) equal to 1 when query nn has a valid box match and 0 otherwise, the discovery-driven contrastive loss is

    Lcontra3D↔Text=∑n=1N1(Fn3DObj,Olabel) CE(Sn,hn),L^{3D\leftrightarrow \mathrm{Text}}_{\mathrm{contra}}=\sum_{n=1}^{N}\mathbf{1}(F^{3D\mathrm{Obj}}_n,O^{\mathrm{label}})\,\mathrm{CE}(S_n,h_n),

    where CE\mathrm{CE} is cross-entropy. For a discovered novel box, the category used to form hnh_n is obtained from the highest CLIP text similarity. Thus, the text alignment learns discriminative point-cloud features over a large vocabulary while restricting category supervision to foreground boxes.

  5. Knowl 5 — Collaborative learning of discovery and alignment

    model/method

    3D-NOD and DCMA are trained as a feedback loop rather than as independent stages. The current class-agnostic detector supplies geometry-based candidate boxes; CLIP image–text similarity filters those candidates and enlarges the novel-box label pool. DCMA then uses the accumulated annotated and discovered boxes to align 3D features with image and text features. The resulting more discriminative 3D object-query features improve objectness and box predictions in subsequent 3D-NOD updates, which supplies additional foreground boxes for alignment. This collaboration simultaneously addresses novel-object localization and novel-category classification: 3D-NOD expands foreground coverage, while DCMA improves the feature representation used to discover and classify those objects.

  6. Knowl 6 — Datasets, category split, and training protocol

    experimental setup

    The method is evaluated on SUN-RGBD and ScanNetV2 under an open-vocabulary split. SUN-RGBD has 5,000 training samples and 46 labeled object categories; the 10 categories with the most training samples are treated as seen and the remaining 36 as novel. ScanNetV2 has 1,200 training samples and 200 object categories; the 10 most frequent categories are treated as seen, while the evaluation uses 50 additional categories as novel. Performance is measured on the validation set using mean average precision and average recall at 3D IoU threshold 0.250.25, denoted by AP25 and AR25; results are also reported separately for novel, base, and all categories.

    CoDA uses 128 object queries and batch size 8. A base 3DETR model is trained for 1,080 epochs with class-agnostic distillation, followed by 200 additional epochs using 3D-NOD and DCMA. The novel-box pool is updated every 50 epochs. Because scenes contain relatively few ground-truth or discovered boxes, the alignment additionally uses 32 object queries selected from the 128 queries beyond those matched to ground-truth boxes. At inference, CoDA uses only point clouds, unlike the image-and-point-cloud comparison pipeline.

  7. Knowl 7 — Component ablations on SUN-RGBD

    data/table

    The component ablations use SUN-RGBD and report AP and AR at IoU 0.250.25 for novel, base, and all categories. The comparison isolates class-agnostic distillation, 3D-NOD, plain text alignment, and the full discovery-driven alignment.

    Could not parse LaTeX table

    Class-agnostic distillation provides preliminary open-vocabulary ability without image input at test time. Adding 3D-NOD raises novel AP from 3.28 to 5.48 and novel AR from 19.87 to 33.45 relative to Distillation. Plain alignment using only base-category text improves base AP but reduces novel AP to 0.85, whereas the discovery-driven full DCMA obtains the best novel and base AP in this comparison.

    A separate collaboration ablation compares applying 3D-NOD without DCMA against the complete jointly learned system:

    Could not parse LaTeX table

    Relative to applying 3D-NOD without DCMA, joint 3D-NOD and DCMA increase novel AP from 5.30 to 6.71 and base AP from 26.08 to 38.72, while also improving novel and base recall. This supports the claimed feedback from feature alignment to object discovery.

  8. Knowl 8 — Comparison with point-cloud and image-based alternatives

    data/table

    The following comparison evaluates open-vocabulary detection on SUN-RGBD and ScanNetV2 at IoU 0.250.25. Det-PointCLIP, Det-PointCLIPv2, and Det-CLIP2 use 3DETR-generated pseudo-boxes and point-cloud open-vocabulary classifiers. 3D-CLIP classifies 3DETR-generated boxes using cropped color-image regions and therefore requires both images and point clouds at test time. CoDA uses only point clouds at test time.

    Could not parse LaTeX table

    CoDA achieves the highest novel AP, mean AP, novel AR, and mean AR on both datasets, while also obtaining the highest base AP and base AR. Its advantage is obtained without requiring color images during testing.

  9. Knowl 9 — Same-setting comparison with OV-3DET

    data/table

    Because OV-3DET uses a different protocol—an external large-scale open-vocabulary 2D detector to generate pseudo-labels and evaluation on only 20 ScanNet categories—the authors also retrained CoDA under that method's setting for a fair comparison. The following values are AP scores for the 20 categories at IoU 0.250.25; Mean is the average over all 20 categories.

    Could not parse LaTeX table
    Could not parse LaTeX table

    CoDA obtains mean AP 19.32 versus 18.02 for OV-3DET, a 1.30-point improvement under the matched protocol, despite the different original training and evaluation settings.

  10. Knowl 10 — Robustness to 3D-NOD confidence thresholds

    data/table

    The 3D-NOD acceptance rule uses a semantic threshold for CLIP category confidence and a geometry threshold for 3D objectness. SUN-RGBD results at IoU 0.250.25 show the effect of varying them independently.

    Could not parse LaTeX table

    The 0.0,0.00.0,0.0 setting is the base Distillation model without 3D-NOD. Every nonzero threshold setting substantially improves novel AP and AR over that baseline. The 0.3,0.30.3,0.3 setting gives the highest novel AP, mean AP, novel AR, and mean AR among the tested settings, while the relatively stable performance across the other settings indicates that 3D-NOD is not dependent on a narrowly tuned threshold pair.

Coverage note — The qualitative scene visualizations and intermediate training-curve plot were omitted because they are illustrative and do not add standalone measurements beyond the reported quantitative ablations and comparisons.

References

  1. 1.Open-set 3d detection via image-level class and debiased cross-modal contrastive learning.
  2. 2.Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. End-to-end object detection with transformers. In ECCV. Springer, 2020.
  3. 3.Bowen Cheng, Lu Sheng, Shaoshuai Shi, Ming Yang, and Dong Xu. Back-tracing representative points for voting-based 3d object detection in point clouds. In CVPR, 2021.
  4. 4.Angela Dai, Angel X Chang, Manolis Savva, Maciej Halber, Thomas Funkhouser, and Matthias Nießner. Scannet: Richly-annotated 3d reconstructions of indoor scenes. In CVPR, 2017.
  5. 5.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In CVPR. IEEE, 2009.
  6. 6.Yu Du, Fangyun Wei, Zihe Zhang, Miaojing Shi, Yue Gao, and Guoqi Li. Learning to prompt for open-vocabulary object detection with vision-language model. In CVPR, 2022.
  7. 7.Chengjian Feng, Yujie Zhong, Zequn Jie, Xiangxiang Chu, Haibing Ren, Xiaolin Wei, Weidi Xie, and Lin Ma. Promptdet: Towards open-vocabulary detection using uncurated images. In ECCV. Springer, 2022.
  8. 8.Xiuye Gu, Tsung-Yi Lin, Weicheng Kuo, and Yin Cui. Open-vocabulary object detection via vision and language knowledge distillation. arXiv preprint arXiv:2104.13921, 2021.
  9. 9.Agrim Gupta, Piotr Dollar, and Ross Girshick. Lvis: A dataset for large vocabulary instance segmentation. In CVPR, 2019.
  10. 10.Akshita Gupta, Sanath Narayan, KJ Joseph, Salman Khan, Fahad Shahbaz Khan, and Mubarak Shah. Ow-detr: Open-world detection transformer. In CVPR, 2022.
  11. 11.Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig. Scaling up visual and vision-language representation learning with noisy text supervision. In ICML. PMLR, 2021.
  12. 12.Aishwarya Kamath, Mannat Singh, Yann LeCun, Gabriel Synnaeve, Ishan Misra, and Nicolas Carion. Mdetr-modulated detection for end-to-end multi-modal understanding. In CVPR, 2021.
  13. 13.Liunian Harold Li, Pengchuan Zhang, Haotian Zhang, Jianwei Yang, Chunyuan Li, Yiwu Zhong, Lijuan Wang, Lu Yuan, Lei Zhang, Jenq-Neng Hwang, et al. Grounded language-image pre-training. In CVPR, 2022.
  14. 14.Ze Liu, Zheng Zhang, Yue Cao, Han Hu, and Xin Tong. Group-free 3d object detection via transformers. In ICCV, 2021.
  15. 15.Yuheng Lu, Chenfeng Xu, Xiaobao Wei, Xiaodong Xie, Masayoshi Tomizuka, Kurt Keutzer, and Shanghang Zhang. Open-vocabulary point-cloud object detection without 3d annotation. arXiv preprint arXiv:2304.00788, 2023.
  16. 16.Xinzhu Ma, Yinmin Zhang, Dan Xu, Dongzhan Zhou, Shuai Yi, Haojie Li, and Wanli Ouyang. Delving into localization errors for monocular 3d object detection. In CVPR, 2021.
  17. 17.Zongyang Ma, Guan Luo, Jin Gao, Liang Li, Yuxin Chen, Shaoru Wang, Congxuan Zhang, and Weiming Hu. Open-vocabulary one-stage detection with hierarchical visual-language knowledge distillation. In CVPR, 2022.
  18. 18.Matthias Minderer, Alexey Gritsenko, Austin Stone, Maxim Neumann, Dirk Weissenborn, Alexey Dosovitskiy, Aravindh Mahendran, Anurag Arnab, Mostafa Dehghani, Zhuoran Shen, et al. Simple open-vocabulary object detection with vision transformers. arXiv preprint arXiv:2205.06230, 2022.
  19. 19.Ishan Misra, Rohit Girdhar, and Armand Joulin. An end-to-end transformer model for 3d object detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 2906–2917, 2021.
  20. 20.Charles R Qi, Hao Su, Kaichun Mo, and Leonidas J Guibas. Pointnet: Deep learning on point sets for 3d classification and segmentation. In CVPR, 2017.
  21. 21.Charles R Qi, Or Litany, Kaiming He, and Leonidas J Guibas. Deep hough voting for 3d object detection in point clouds. In ICCV, 2019.
  22. 22.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In ICML. PMLR, 2021.
  23. 23.David Rozenberszki, Or Litany, and Angela Dai. Language-grounded indoor 3d semantic segmentation in the wild. In ECCV. Springer, 2022.
  24. 24.Shuran Song, Samuel P Lichtenberg, and Jianxiong Xiao. Sun rgb-d: A rgb-d scene understanding benchmark suite. In CVPR, 2015.
  25. 25.Qian Xie, Yu-Kun Lai, Jing Wu, Zhoutao Wang, Yiming Zhang, Kai Xu, and Jun Wang. Mlcvnet: Multi-level context votenet for 3d object detection. In CVPR, 2020.
  26. 26.Qian Xie, Yu-Kun Lai, Jing Wu, Zhoutao Wang, Dening Lu, Mingqiang Wei, and Jun Wang. Venet: Voting enhancement network for 3d object detection. In ICCV, 2021.
  27. 27.Lewei Yao, Jianhua Han, Youpeng Wen, Xiaodan Liang, Dan Xu, Wei Zhang, Zhenguo Li, Chunjing Xu, and Hang Xu. Detclip: Dictionary-enriched visual-concept paralleled pre-training for open-world detection. In NeurIPS, 2022.
  28. 28.Lewei Yao, Jianhua Han, Xiaodan Liang, Dan Xu, Wei Zhang, Zhenguo Li, and Hang Xu. Detclipv2: Scalable open-vocabulary object detection pre-training via word-region alignment. In CVPR, 2023.
  29. 29.Alireza Zareian, Kevin Dela Rosa, Derek Hao Hu, and Shih-Fu Chang. Open-vocabulary object detection using captions. In CVPR, 2021.
  30. 30.Yihan Zeng, Chenhan Jiang, Jiageng Mao, Jianhua Han, Chaoqiang Ye, Qingqiu Huang, Dit-Yan Yeung, Zhen Yang, Xiaodan Liang, and Hang Xu. Clip2: Contrastive language-image-point pretraining from real-world point cloud data. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 15244–15253, 2023.
  31. 31.Renrui Zhang, Ziyu Guo, Wei Zhang, Kunchang Li, Xupeng Miao, Bin Cui, Yu Qiao, Peng Gao, and Hongsheng Li. Pointclip: Point cloud understanding by clip. In CVPR, 2022.
  32. 32.Yinmin Zhang, Xinzhu Ma, Shuai Yi, Jun Hou, Zhihui Wang, Wanli Ouyang, and Dan Xu. Learning geometry-guided depth via projective modeling for monocular 3d object detection. arXiv preprint arXiv:2107.13931, 2021.
  33. 33.Zaiwei Zhang, Bo Sun, Haitao Yang, and Qixing Huang. H3dnet: 3d object detection using hybrid geometric primitives. In ECCV. Springer, 2020.
  34. 34.Yiwu Zhong, Jianwei Yang, Pengchuan Zhang, Chunyuan Li, Noel Codella, Liunian Harold Li, Luowei Zhou, Xiyang Dai, Lu Yuan, Yin Li, et al. Regionclip: Region-based language-image pretraining. In CVPR, 2022.
  35. 35.Xingyi Zhou, Rohit Girdhar, Armand Joulin, Phillip Krähenbühl, and Ishan Misra. Detecting twenty-thousand classes using image-level supervision. arXiv preprint arXiv:2201.02605, 2022.
  36. 36.Xiangyang Zhu, Renrui Zhang, Bowei He, Ziyao Zeng, Shanghang Zhang, and Peng Gao. Pointclip v2: Adapting clip for powerful 3d open-world learning. arXiv preprint arXiv:2211.11682, 2022.

Citation

MLA
Cao, Y., et al. “CoDA: Collaborative Novel Box Discovery and Cross-modal Alignment for Open-vocabulary 3D Object Detection”. arXiv, 2023, http://arxiv.org/abs/2310.02960v1.
APA
Cao, Y., Zeng, Y., Xu, H., & Xu, D. (2023). CoDA: Collaborative Novel Box Discovery and Cross-modal Alignment for Open-vocabulary 3D Object Detection. arXiv. http://arxiv.org/abs/2310.02960v1
Chicago
Cao, Y., Y. Zeng, H. Xu, and D. Xu. 2023. “CoDA: Collaborative Novel Box Discovery and Cross-modal Alignment for Open-vocabulary 3D Object Detection”. arXiv. http://arxiv.org/abs/2310.02960v1.
Harvard
Cao, Y. et al. (2023) “CoDA: Collaborative Novel Box Discovery and Cross-modal Alignment for Open-vocabulary 3D Object Detection”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2310.02960v1.
Vancouver
1. Cao Y, Zeng Y, Xu H, Xu D (2023) CoDA: Collaborative Novel Box Discovery and Cross-modal Alignment for Open-vocabulary 3D Object Detection. arXiv

BibTeX

@article{cao2023coda,
  title = {CoDA: Collaborative Novel Box Discovery and Cross-modal Alignment for Open-vocabulary 3D Object Detection},
  author = {Cao, Yang and Zeng, Yihan and Xu, Hang and Xu, Dan},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2310.02960v1},
  eprint = {2310.02960}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors