CoDA: Collaborative Novel Box Discovery and Cross-modal Alignment for Open-vocabulary 3D Object Detection
Yang CaoYihan ZengHang XuDan Xu
Proposes an end-to-end framework that couples 3D novel box discovery with cross-modal alignment to simultaneously localize and classify unseen 3D objects without relying on external 2D open-vocabulary detectors.
Three-dimensional (3D) object detection plays a vital role in real-world technologies such as autonomous driving, robotics, and industrial manufacturing. However, traditional systems are restricted to fixed, pre-defined sets of categories because annotating 3D point cloud data is extremely labor-intensive and costly. Open-vocabulary 3D object detection addresses this limitation by enabling models to recognize and localize novel, unbounded categories. The core challenge lies in simultaneously identifying the location of unseen objects and classifying them accurately when only a very small set of base categories has human annotations.
The main objective of the article is to demonstrate a unified framework, named CoDA, that simultaneously discovers novel 3D bounding boxes and classifies unseen objects without relying on complex, external two-dimensional (2D) open-vocabulary detection models. It evaluates the performance of this framework against existing baseline and alternative detection methods across benchmark indoor datasets.
To achieve this, the article introduces a collaborative approach combining a 3D Novel Object Discovery strategy with a Discovery-Driven Cross-Modal Alignment module. The novel object discovery strategy leverages 3D geometric bounding box patterns from known base categories alongside 2D semantic priors from the pre-trained CLIP vision-language model to propose and verify pseudo bounding boxes for unseen objects. The cross-modal alignment module then aligns 3D point cloud features with 2D image and text representations through category-agnostic distillation and class-specific contrastive learning. The discovery of novel boxes and the alignment of multi-modal features are trained iteratively, where improvements in one directly enhance the other. The framework was evaluated on two widely used benchmark datasets: SUN-RGBD (using 10 base and 36 novel categories) and ScanNet (using 10 base and 50 novel categories).
The findings show substantial improvements over previous techniques. First, the unified CoDA framework outperformed the best-performing alternative methods by more than 80% in mean Average Precision (mAP) for 3D object detection. Second, the novel object discovery mechanism increased novel category recall from approximately 21.5% in the baseline to 33.7% in the complete framework on the SUN-RGBD dataset, overcoming category forgetting during extended training. Third, collaborative learning between object discovery and cross-modal feature alignment improved precision on novel categories by roughly 26% and base categories by 48% compared to discovery alone. Fourth, during deployment and testing, the framework successfully detected and categorized novel objects using purely 3D point cloud data without requiring 2D camera images.
These results demonstrate that systems can successfully adapt to open-world environments with minimal manual annotation, significantly reducing the cost and timeline associated with data labeling. Unlike prior approaches that require heavy secondary 2D detectors, the proposed method provides an end-to-end pathway that reduces system dependencies while maintaining high detection accuracy. For robotics and autonomous systems, this expands operational safety and flexibility in novel, unstructured environments.
Organizations developing spatial computing, robotics, or autonomous vehicle systems should consider incorporating collaborative 3D-to-2D feature alignment architectures to handle unknown object classes. Technical teams should run pilot evaluations on their domain-specific point cloud data to determine if discovery-driven distillation can reduce manual labeling overhead. Because current evaluations remain focused on indoor benchmark environments, further validation on outdoor datasets and under noisy sensor conditions is recommended before full-scale commercial deployment.
- Paper: Open-vocabulary Object Detection via Vision and Language Knowledge Distillation, Xiuye Gu et al. (2021). ViLD establishes the foundational paradigm of distilling vision-and-language knowledge from pre-trained teachers into object detectors, which CoDA directly adapts and extends to 3D point cloud object discovery and cross-modal alignment.
- Paper: Deep Hough Voting for 3D Object Detection in Point Clouds, Charles R. Qi et al. (2019). VoteNet provides the core 3D indoor object proposal and bounding box regression architecture that serves as the baseline detector and geometric foundation for CoDA.
- Paper: OW-DETR: Open-world Detection Transformer, Akshita Gupta et al. (2022). OW-DETR formalizes open-world pseudo-labeling and objectness scoring to discover novel unknown classes, establishing key principles underpinning CoDA's 3D novel object discovery.
- Paper: PointNet: Deep Learning on Point Sets for 3D Classification and Segmentation, Charles R. Qi et al. (2017). PointNet introduces the fundamental neural architecture for processing raw, unstructured 3D point cloud sets utilized across modern 3D recognition backbones.
- Paper: WildDet3D: Scaling Promptable 3D Detection in the Wild, Weikai Huang et al. (2026). WildDet3D scales promptable, open-vocabulary 3D detection to in-the-wild settings across thousands of categories, generalizing beyond the indoor benchmark discoveries developed in CoDA.
