Language Conditioned Spatial Relation Reasoning for 3D Object Grounding

Shizhe ChenPierre-Louis GuhurMakarand TapaswiCordelia SchmidIvan Laptev

article2022NeurIPS149 citations

Proposes a transformer architecture with language-conditioned spatial self-attention and teacher-student knowledge distillation to improve 3D object grounding by accurately reasoning over relative distances and orientations in point clouds.

Listen

Enabling autonomous systems and robots to accurately locate physical items in three-dimensional environments from natural language commands is a foundational challenge in modern artificial intelligence. In practical settings, language frequently distinguishes between identical objects using spatial relationships, such as identifying the nearest backpack or selecting a door to the left. However, standard machine learning architectures often fail to resolve these relationships effectively in raw 3D point cloud data, and training is severely constrained by the scarcity of annotated 3D data compared to 2D image domains.

The article evaluates and demonstrates a new vision-and-language framework designed to ground 3D objects and reason about their spatial configurations directly from natural language descriptions. Specifically, it tests whether integrating explicit geometric relationships into self-attention layers alongside a targeted knowledge distillation strategy can significantly outperform existing state-of-the-art approaches.

The researchers developed the ViL3DRel model, which incorporates a specialized spatial self-attention layer into a multimodal transformer architecture to encode pairwise relative distances and horizontal and vertical orientations between objects. To address the problem of noisy 3D point cloud representations during training, the team introduced a teacher-student framework without relying on external 2D images. A teacher model was first trained using ground-truth object labels and dominant colors to master relational reasoning, after which its intermediate attention and hidden representations were distilled into a student model operating solely on raw point cloud inputs. The approach was evaluated across three widely recognized benchmark datasets based on real-world indoor scans: Nr3D, Sr3D, and ScanRefer.

The evaluation produced several decisive findings. First, the proposed framework achieved substantial performance gains over previous state-of-the-art models, reaching 64.4% accuracy on Nr3D and 72.8% on Sr3D given ground-truth object proposals, representing absolute improvements of 9.3 and 8.3 percentage points respectively. Second, on the ScanRefer dataset using automatically detected proposals, the model achieved an overall accuracy of 37.73% at the strict 0.5 intersection-over-union metric, outperforming the prior benchmark of 33.26%. Third, ablation experiments revealed that combining explicit distance and orientation modeling is crucial: distance features improved view-independent queries, while orientation features drove a 10.2 percentage point gain on view-dependent queries. Finally, transferring intermediate attention weights and hidden states from the teacher model provided a clear performance boost of over 6 percentage points compared to training the student model from scratch.

These findings demonstrate that spatial reasoning in complex environments requires explicit geometric structure rather than relying purely on implicit learning within standard transformer layers. Operationally, the teacher-student strategy provides a highly cost-effective training paradigm by bypassing the requirement for costly paired 2D imagery or camera calibrations. Improving 3D grounding accuracy directly enhances the reliability, safety, and autonomy of robotic assistants navigating cluttered physical environments.

Organizations developing embodied AI and robotics should adopt explicit spatial attention mechanisms and cross-modal distillation strategies to enhance spatial comprehension. Further research should focus on integrating these spatial reasoning layers directly into single-stage, end-to-end 3D object detection pipelines to eliminate bottlenecks caused by imperfect initial proposal generation. Leaders should note that confidence in these results is supported by rigorous multi-benchmark evaluations, though practical deployment cautions remain due to performance degradation when relying on noisy automated object detectors rather than ground-truth proposals, as well as the limited environmental diversity present in existing benchmark datasets.

Cover for Language Conditioned Spatial Relation Reasoning for 3D Object Grounding

Abstract

Localizing objects in 3D scenes based on natural language requires understanding and reasoning about spatial relations. In particular, it is often crucial to distinguish similar objects referred by the text, such as "the left most chair" and "a chair next to the window". In this work we propose a language-conditioned transformer model for grounding 3D objects and their spatial relations. To this end, we design a spatial self-attention layer that accounts for relative distances and orientations between objects in input 3D point clouds. Training such a layer with visual and language inputs enables to disambiguate spatial relations and to localize objects referred by the text. To facilitate the cross-modal learning of relations, we further propose a teacher-student approach where the teacher model is first trained using ground-truth object labels, and then helps to train a student model using point cloud inputs. We perform ablation studies showing advantages of our approach. We also demonstrate our model to significantly outperform the state of the art on the challenging Nr3D, Sr3D and ScanRefer 3D object grounding datasets.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Method
  • 3.1 Architecture Overview
  • 3.2 Spatial Self-Attention
  • 3.3 Teacher-Student Training
  • 4 Experiments
  • 4.1 Datasets
  • 4.2 Experimental Setting
  • 4.3 Ablation Studies
  • 4.3.1 Spatial Relation Reasoning
  • 4.3.2 Teacher-Student Training
  • 4.4 Comparison with State-of-the-Art Methods
  • 5 Conclusion
  • Acknowledgments and Disclosure of Funding
  • References
  • Checklist

Knowls

  1. Knowl 1 — ViL3DRel Architecture for 3D Object Grounding

    model/method

    The Vision-and-Language 3D Relation reasoning model (ViL3DRel) addresses the 3D object grounding task: identifying and selecting a referred target 3D bounding box BT∈R6B_T \in \mathbb{R}^6 in a 3D scene point cloud Pscene∈RK×6\mathcal{P}_{scene} \in \mathbb{R}^{K \times 6} given a natural language sentence SS. Operating within a two-stage detection-then-matching framework, the architecture assumes NN candidate object proposals (O1,…,ON)(O_1, \dots, O_N), where each proposal OiO_i contains a subset of points Pi∈RKi×6P_i \in \mathbb{R}^{K_i \times 6}.

    The framework consists of four main modules:

    1. Text Encoding: A 3-layer transformer initialized from BERT encodes the MM-word sentence SS into token features (scls,s1,…,sM)(s_{cls}, s_1, \dots, s_M), where scls,sj∈Rds_{cls}, s_j \in \mathbb{R}^d (d=768d=768).

    2. Object Encoding: For each proposal OiO_i, point coordinates in PiP_i are normalized into a unit ball and processed by a fixed PointNet++ network to extract visual point cloud feature oi0∈Rdo_i^0 \in \mathbb{R}^d. The proposal's 3D bounding center ci=[cx,cy,cz]∈R3c_i = [c_x, c_y, c_z] \in \mathbb{R}^3 (mean of PiP_i) and 3D spatial extent zi=[zx,zy,zz]∈R3z_i = [z_x, z_y, z_z] \in \mathbb{R}^3 are projected into an absolute 3D position embedding li=Wl[ci;zi]∈Rdl_i = W_l [c_i; z_i] \in \mathbb{R}^d via a linear layer WlW_l.

    3. Multimodal Fusion: A stack of L=4L=4 transformer layers fuses textual and visual features. In layer ll, the input visual token is augmented with its absolute location feature (oil+lio_i^l + l_i), processed by a language-conditioned spatial self-attention layer to yield contextualized representation o^il\hat{o}_i^l, passed to a cross-attention layer that uses o^il\hat{o}_i^l as queries and text representations as keys/values, and finalized by a 2-layer feed-forward network (FFN).

    4. Grounding Head: A two-layer FFN scores each multimodal proposal token oiLo_i^L from the final fusion layer and applies a softmax over all NN proposals to predict proposal selection probabilities pip_i. The proposal with the highest probability is selected as the referred target.

  2. Knowl 2 — Language-Conditioned Spatial Self-Attention Mechanism

    model/method

    To disambiguate same-class objects referred to by relative spatial descriptions, ViL3DRel incorporates an explicit spatial self-attention layer into the multimodal transformer. For NN object proposals, standard self-attention computes query, key, and value matrices Q=XWQ,K=XWK,V=XWV∈RN×dhQ = X W_Q, K = X W_K, V = X W_V \in \mathbb{R}^{N \times d_h} from object features X∈RN×dX \in \mathbb{R}^{N \times d}, producing appearance attention logits ωijo=QiKjTdh\omega_{ij}^o = \frac{Q_i K_j^T}{\sqrt{d_h}}.

    To explicitly account for relative distances and relative orientations between objects (Oi,Oj)(O_i, O_j):

    1. An explicit 5-dimensional pairwise spatial feature vector fijs∈R5f_{ij}^s \in \mathbb{R}^5 is computed: fijs=[dij,sin⁡(θh),cos⁡(θh),sin⁡(θv),cos⁡(θv)]f_{ij}^s = [d_{ij}, \sin(\theta_h), \cos(\theta_h), \sin(\theta_v), \cos(\theta_v)] where dij=∥ci−cj∥2d_{ij} = \|c_i - c_j\|_2 is the Euclidean distance between 3D proposal centers ci,cj∈R3c_i, c_j \in \mathbb{R}^3, and θh,θv\theta_h, \theta_v are horizontal and vertical angles of the vector connecting cic_i to cjc_j.

    2. A language-conditioned gating weight gis∈R5g_i^s \in \mathbb{R}^5 is predicted for proposal OiO_i to select relevant spatial relations conditioned on the sentence: gis=WST(scls+oil)g_i^s = W_S^T (s_{cls} + o_i^l) where WS∈Rd×5W_S \in \mathbb{R}^{d \times 5} is a learnable projection matrix, sclss_{cls} is the sentence CLS token, and oilo_i^l is the object feature at layer ll.

    3. The spatial relevance score between objects OiO_i and OjO_j is given by: ωijs=gis⋅fijs\omega_{ij}^s = g_i^s \cdot f_{ij}^s

    Multi-head attention across H=12H=12 heads is employed, allowing independent spatial relation reasoning across heads before concatenation.

  3. Knowl 3 — Sigmoid-Softmax Attention Fusion Formulation

    equation

    In ViL3DRel's spatial self-attention module, the language-conditioned pairwise spatial relevance ωijs\omega_{ij}^s is fused with the standard scaled dot-product visual attention logit ωijo=QiKjTdh\omega_{ij}^o = \frac{Q_i K_j^T}{\sqrt{d_h}} for object proposal query ii and key jj (i,j∈{1,…,N}i, j \in \{1, \dots, N\}) via a sigmoid-softmax (sigsoftmax) modulation function:

    ωij=σ(ωijs)exp⁡(ωijo)∑l=1Nσ(ωils)exp⁡(ωilo)\omega_{ij} = \frac{\sigma(\omega_{ij}^s) \exp(\omega_{ij}^o)}{\sum_{l=1}^N \sigma(\omega_{il}^s) \exp(\omega_{il}^o)}

    where σ(x)=11+exp⁡(−x)\sigma(x) = \frac{1}{1 + \exp(-x)} is the sigmoid activation function, and Ω=[ωij]N×N\Omega = [\omega_{ij}]_{N \times N} forms the normalized attention weight matrix used to aggregate the value vectors VV as SpatialSelfAttn(Q,K,V)=ΩV\text{SpatialSelfAttn}(Q, K, V) = \Omega V.

  4. Knowl 4 — Teacher-Student Knowledge Distillation for 3D Grounding

    model/method

    To overcome performance degradation caused by noisy 3D point cloud object representations, ViL3DRel uses a teacher-student training framework. The teacher and student share the same transformer architecture but differ in their object inputs:

    • Teacher Object Representation: Uses ground-truth object category labels and dominant colors. Ground-truth class labels are embedded via pre-trained GloVe word vectors. Dominant colors are extracted by fitting a 3-component Gaussian Mixture Model (GMM) on point RGB values for each object, projecting each component's mean RGB vector via a linear layer, and computing their mixture-weighted sum. The teacher's initial object representation is the sum of the class embedding and the aggregated color embedding.
    • Student Object Representation: Uses object visual features extracted from 3D point clouds via PointNet++.

    Knowledge is distilled from the pre-trained teacher to the student model via two distillation loss functions:

    1. Attention Distillation Loss (Lattn\mathcal{L}_{attn}): Enforces the student to mimic all self-attention and cross-attention matrices of the teacher across all L=4L=4 layers and H=12H=12 heads: Lattn=1LH∑l=1L∑h=1HMSE(ΩlhS−ΩlhT)\mathcal{L}_{attn} = \frac{1}{LH} \sum_{l=1}^L \sum_{h=1}^H \text{MSE}(\Omega_{lh}^S - \Omega_{lh}^T) where ΩlhS\Omega_{lh}^S and ΩlhT\Omega_{lh}^T are student and teacher attention matrices at layer ll, head hh.

    2. Hidden State Distillation Loss (Lhidden\mathcal{L}_{hidden}): Distills output representations across all layers l∈{0,…,L}l \in \{0, \dots, L\} for all NN proposals: Lhidden=1LN∑l=0L∑i=1NMSE(oil,S−oil,T)\mathcal{L}_{hidden} = \frac{1}{LN} \sum_{l=0}^L \sum_{i=1}^N \text{MSE}(o_i^{l,S} - o_i^{l,T}) where oil,S,oil,T∈Rdo_i^{l,S}, o_i^{l,T} \in \mathbb{R}^d are student and teacher token embeddings for object proposal ii at layer ll.

  5. Knowl 5 — Multi-Loss Training Objective for ViL3DRel

    equation

    The overall training objective L\mathcal{L} for training the ViL3DRel student model combines object grounding loss, auxiliary classification losses, and distillation losses:

    L=Log+Lsent+Lobju+Lobjm+λaLattn+λhLhidden\mathcal{L} = \mathcal{L}_{og} + \mathcal{L}_{sent} + \mathcal{L}_{obj}^u + \mathcal{L}_{obj}^m + \lambda_a \mathcal{L}_{attn} + \lambda_h \mathcal{L}_{hidden}

    where:

    • Log\mathcal{L}_{og} is the cross-entropy 3D object grounding loss over proposal predictions pip_i.
    • Lsent\mathcal{L}_{sent} is a cross-entropy sentence classification loss predicting the target category from the text classification token sclss_{cls}.
    • Lobju\mathcal{L}_{obj}^u is an object category classification loss based on unimodal visual proposal features oi0o_i^0.
    • Lobjm\mathcal{L}_{obj}^m is an object category classification loss based on multimodal fused proposal features oiLo_i^L.
    • Lattn\mathcal{L}_{attn} is the mean squared error attention distillation loss.
    • Lhidden\mathcal{L}_{hidden} is the mean squared error hidden state distillation loss.
    • Distillation loss hyperparameters are set to λa=1.0\lambda_a = 1.0 and λh=0.02\lambda_h = 0.02.
  6. Knowl 6 — Grounding Accuracy on Nr3D and Sr3D Benchmarks

    data/table

    Evaluated using ground-truth 3D object proposals on the human-annotated Nr3D dataset and template-generated Sr3D dataset, ViL3DRel outperforms existing methods across all evaluation splits (overall, easy distractors, hard distractors, viewpoint-dependent, and viewpoint-independent descriptions):

    Method Nr3D Sr3D
    Overall Easy Hard ViewDep ViewIndep Overall Easy Hard ViewDep ViewIndep
    ReferIt3D 35.6 43.6 27.9 32.5 37.1 40.8 44.7 31.5 39.2 40.8
    ScanRefer 34.2 41.0 23.5 29.9 35.4 - - - - -
    TGNN 37.3 44.2 30.6 35.8 38.0 - - - - -
    InstanceRefer 38.8 46.0 31.8 34.5 41.9 48.0 51.1 40.5 45.4 48.1
    FFL-3DOG 41.7 48.2 35.0 37.1 44.7 - - - - -
    3DVG-Trans 40.8 48.5 34.8 34.8 43.7 51.4 54.2 44.9 44.6 51.7
    TransRefer3D 42.1 48.5 36.0 36.5 44.9 57.4 60.5 50.2 49.9 57.7
    LanguageRefer 43.9 51.0 36.6 41.7 45.0 56.0 58.9 49.3 49.2 56.3
    SAT 49.2 56.3 42.4 46.9 50.4 57.9 61.2 50.0 49.2 58.3
    3D-SPS 51.5 58.1 45.1 48.0 53.2 62.6 56.2 65.4 49.2 63.2
    Multi-view 55.1 61.3 49.1 54.3 55.4 64.5 66.9 58.8 58.4 64.7
    ViL3DRel (Ours) 64.4 70.2 57.4 62.0 64.5 72.8 74.9 67.9 63.8 73.2

    Compared to the previous best approach (Multi-view), ViL3DRel achieves an absolute gain of +9.3%+9.3\% on Nr3D (from 55.1%55.1\% to 64.4%64.4\%) and +8.3%+8.3\% on Sr3D (from 64.5%64.5\% to 72.8%72.8\%).

  7. Knowl 7 — Grounding Accuracy on ScanRefer Benchmark

    data/table

    On the ScanRefer dataset, performance is evaluated with ground-truth proposals and detected object proposals (using pre-trained PointGroup object detections). Metrics are bounding box IoU accuracies at thresholds 0.25 ([email protected]) and 0.5 ([email protected]) across "Unique" (single object of category in scene), "Multiple" (distractors present), and "Overall" categories:

    Method Detector Unique Multiple Overall GT Props
    [email protected] [email protected] [email protected] [email protected] [email protected] [email protected] Overall Acc (%)
    ReferIt3D - - - - - - - 46.9
    Non-SAT VN 68.48 47.38 31.81 21.34 38.92 26.40 48.2
    SAT VN 73.21 50.83 37.64 25.16 44.54 30.14 53.8
    InstanceRefer PG 77.45 66.83 31.27 24.77 40.23 32.93 -
    Multi-view PG 77.67 66.45 31.92 25.26 40.80 33.26 -
    ViL3DRel (Ours) PG 81.58 68.62 40.30 30.71 47.94 37.73 59.8
    Upper Bound PG 88.63 74.47 78.82 60.37 80.64 62.98 100.0

    (Note: VN = VoteNet, PG = PointGroup).

    Using identical PointGroup proposals, ViL3DRel outperforms Multi-view by +7.14+7.14 points in [email protected] (from 40.80%40.80\% to 47.94%47.94\%) and +4.47+4.47 points in [email protected] (from 33.26%33.26\% to 37.73%37.73\%). Given ground-truth proposals, ViL3DRel achieves 59.8%59.8\%, exceeding SAT (53.8%53.8\%).

  8. Knowl 8 — Ablation Analysis of Spatial Attention Features and Fusion Modes

    data/table

    Ablations on the Nr3D teacher model (with ground-truth object labels) demonstrate the individual and combined impact of pairwise spatial relation features, attention head structures, attention fusion functions, and data augmentation:

    Row Pairwise Dist Pairwise Ort Multi-Head Spat. Fusion Mode RotAug Overall (%) ViewDep (%) ViewIndep (%)
    R1 - - - - No 53.5 51.4 54.6
    R2 (+Color) - - - - No 55.1 53.8 55.8
    R3 (+RotAug) - - - - Yes 62.4 58.3 64.5
    R4 (Dist only) Yes No Yes sigs Yes 66.0 53.8 72.0
    R5 (Ort only) No Yes Yes sigs Yes 71.3 68.5 72.6
    R6 (Single Head) Yes Yes No sigs Yes 67.7 65.2 69.0
    R7 (Bias RPE) Yes Yes Yes bias Yes 55.4 46.8 59.6
    R8 (Context RPE) Yes Yes Yes ctx Yes 56.4 50.8 59.1
    R9 (Full ViL3DRel) Yes Yes Yes sigs Yes 74.4 71.3 75.9

    Key takeaways include:

    1. Distance vs. Orientation: Distance features primarily improve view-independent grounding (64.5%→72.0%64.5\% \to 72.0\%) without helping view-dependent queries (58.3%→53.8%58.3\% \to 53.8\%). Orientation features boost view-dependent grounding substantially (58.3%→68.5%58.3\% \to 68.5\%). Combining both achieves 74.4%74.4\%.
    2. Attention Fusion: Sigmoid-softmax (sigs, 74.4%74.4\%) markedly outperforms conventional relative positional encoding (RPE) mechanisms such as additive attention bias (bias, 55.4%55.4\%) and contextual projection (ctx, 56.4%56.4\%).
    3. Spatial Feature Construction: Computing geometric features from 3D object centers achieves 74.4%74.4\% (identical to using bottom centers at 74.4%74.4\%), whereas an MLP trained directly on concatenated 3D bounding box coordinates only achieves 57.4%57.4\%.
  9. Knowl 9 — Ablation Analysis of Distillation Objectives in Student Learning

    data/table

    Ablation of knowledge transfer components from the teacher model (using ground-truth semantic object representations) to the student model (using PointNet++ point cloud representations) on the Nr3D dataset:

    Model Setup Teacher Weight Init Lattn\mathcal{L}_{attn} Lhidden\mathcal{L}_{hidden} Grounding Accuracy (%)
    Teacher (GT labels) - - - 74.4
    Student Baseline No No No 58.1
    Student + Weight Init Yes No No 62.6
    Student + Attention Distillation No Yes No 63.6
    Student + Hidden Distillation No No Yes 62.1
    Student + Full Distillation (ViL3DRel) No Yes Yes 64.4

    Distilling intermediate attention maps (Lattn\mathcal{L}_{attn}, +5.5%+5.5\%) transfers relational structure more effectively than distilling hidden states alone (Lhidden\mathcal{L}_{hidden}, +4.0%+4.0\%). Combining Lattn\mathcal{L}_{attn} and Lhidden\mathcal{L}_{hidden} achieves the highest performance (64.4%64.4\%, +6.3%+6.3\% over the undistilled student baseline).

  10. Knowl 10 — Limitations of the ViL3DRel Framework

    limitation

    The ViL3DRel framework exhibits three primary limitations:

    1. Dependence on Two-Stage Proposal Generation: Grounding accuracy is constrained by the recall and quality of the upstream 3D instance segmentation / proposal network. On ScanRefer, grounding accuracy drops by more than 12 percentage points when shifting from ground-truth object proposals (59.8%59.8\%) to automatic PointGroup detector proposals (47.94%47.94\%).
    2. Lack of Explicit Object Orientation and Pose Estimation: Pairwise relative spatial features are calculated between 3D object centroids using global coordinate orientations rather than object-centric reference frames, as predicting accurate 3D poses for arbitrary point cloud objects remains difficult.
    3. Environmental Diversity in Datasets: Evaluations are restricted to ScanNet-derived indoor environments (Nr3D, Sr3D, ScanRefer), which have bounded visual and geometric diversity.

Coverage note — None omitted; all primary methodology, architectural details, loss functions, distillation mechanisms, benchmark evaluations, and ablation studies have been captured.

References

  1. 1.Licheng Yu, Patrick Poirson, Shan Yang, Alexander C Berg, and Tamara L Berg. Modeling context in referring expressions. In ECCV, pages 69–85. Springer, 2016.
  2. 2.Junhua Mao, Jonathan Huang, Alexander Toshev, Oana Camburu, Alan L Yuille, and Kevin Murphy. Generation and comprehension of unambiguous object descriptions. In CVPR, pages 11–20, 2016.
  3. 3.Runtao Liu, Chenxi Liu, Yutong Bai, and Alan L Yuille. Clevr-ref+: Diagnosing visual reasoning with referring expressions. In CVPR, pages 4185–4194, 2019.
  4. 4.Zhenfang Chen, Peng Wang, Lin Ma, Kwan-Yee K Wong, and Qi Wu. Cops-ref: A new dataset and task on compositional referring expression comprehension. In CVPR, pages 10086–10095, 2020.
  5. 5.Howard Chen, Alane Suhr, Dipendra Misra, Noah Snavely, and Yoav Artzi. Touchdown: Natural language navigation and spatial reasoning in visual street environments. In CVPR, pages 12538–12547, 2019.
  6. 6.Yuankai Qi, Qi Wu, Peter Anderson, Xin Wang, William Yang Wang, Chunhua Shen, and Anton van den Hengel. Reverie: Remote embodied visual referring expression in real indoor environments. In CVPR, pages 9982–9991, 2020.
  7. 7.Dave Zhenyu Chen, Angel X Chang, and Matthias Nießner. Scanrefer: 3d object localization in rgb-d scans using natural language. In ECCV, pages 202–221. Springer, 2020.
  8. 8.Panos Achlioptas, Ahmed Abdelreheem, Fei Xia, Mohamed Elhoseiny, and Leonidas Guibas. Referit3d: Neural listeners for fine-grained 3d object identification in real-world scenes. In ECCV, pages 422–440. Springer, 2020.
  9. 9.Dailan He, Yusheng Zhao, Junyu Luo, Tianrui Hui, Shaofei Huang, Aixi Zhang, and Si Liu. Transrefer3d: Entity-and-relation aware transformer for fine-grained 3d visual grounding. In ACM MM, pages 2344–2352, 2021.
  10. 10.Pin-Hao Huang, Han-Hung Lee, Hwann-Tzong Chen, and Tyng-Luh Liu. Text-guided graph neural networks for referring 3d instance segmentation. In AAAI, volume 35, pages 1610–1618, 2021.
  11. 11.Zhihao Yuan, Xu Yan, Yinghong Liao, Ruimao Zhang, Sheng Wang, Zhen Li, and Shuguang Cui. Instancerefer: Cooperative holistic understanding for visual grounding on point clouds through instance multi-level contextual referring. In ICCV, pages 1791–1800, 2021.
  12. 12.Mingtao Feng, Zhen Li, Qi Li, Liang Zhang, XiangDong Zhang, Guangming Zhu, Hui Zhang, Yaonan Wang, and Ajmal Mian. Free-form description guided 3d visual graph network for object grounding in point cloud. In ICCV, pages 3722–3731, 2021.
  13. 13.Lichen Zhao, Daigang Cai, Lu Sheng, and Dong Xu. 3dvg-transformer: Relation modeling for visual grounding on point clouds. In ICCV, pages 2928–2937, 2021.
  14. 14.Zhengyuan Yang, Songyang Zhang, Liwei Wang, and Jiebo Luo. Sat: 2d semantics assisted training for 3d visual grounding. In ICCV, pages 1856–1866, 2021.
  15. 15.Junha Roh, Karthik Desingh, Ali Farhadi, and Dieter Fox. Languagerefer: Spatial-language model for 3d visual grounding. In CoRL, pages 1046–1056. PMLR, 2021.
  16. 16.Shijia Huang, Yilun Chen, Jiaya Jia, and Liwei Wang. Multi-view transformer for 3d visual grounding. In CVPR, 2022.
  17. 17.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. NeurIPS, 30, 2017.
  18. 18.Aishwarya Kamath, Mannat Singh, Yann LeCun, Gabriel Synnaeve, Ishan Misra, and Nicolas Carion. Mdetr-modulated detection for end-to-end multi-modal understanding. In ICCV, pages 1780–1790, 2021.
  19. 19.Sanjay Subramanian, Will Merrill, Trevor Darrell, Matt Gardner, Sameer Singh, and Anna Rohrbach. Reclip: A strong zero-shot baseline for referring expression comprehension. In ACL, 2022.
  20. 20.Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A Shamma, et al. Visual genome: Connecting language and vision using crowdsourced dense image annotations. IJCV, 123(1):32–73, 2017.
  21. 21.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In ICML, pages 8748–8763. PMLR, 2021.
  22. 22.Project webpage. https://cshizhe.github.io/projects/vil3dref.html.
  23. 23.Junyu Luo, Jiahui Fu, Xianghao Kong, Chen Gao, Haibing Ren, Hao Shen, Huaxia Xia, and Si Liu. 3d-sps: Single-stage 3d visual grounding via referred point progressive selection. In CVPR, 2022.
  24. 24.Licheng Yu, Zhe Lin, Xiaohui Shen, Jimei Yang, Xin Lu, Mohit Bansal, and Tamara L Berg. Mattnet: Modular attention network for referring expression comprehension. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 1307–1315, 2018.
  25. 25.Angela Dai, Angel X Chang, Manolis Savva, Maciej Halber, Thomas Funkhouser, and Matthias Nießner. Scannet: Richly-annotated 3d reconstructions of indoor scenes. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 5828–5839, 2017.
  26. 26.Petar Veličković, Guillem Cucurull, Arantxa Casanova, Adriana Romero, Pietro Liò, and Yoshua Bengio. Graph attention networks. ICLR, 2018.
  27. 27.Yue Wang, Yongbin Sun, Ziwei Liu, Sanjay E Sarma, Michael M Bronstein, and Justin M Solomon. Dynamic graph cnn for learning on point clouds. TOG, 38(5):1–12, 2019.
  28. 28.Ayush Jain, Nikolaos Gkanatsios, Ishita Mediratta, and Katerina Fragkiadaki. Looking outside the box to ground language in 3d scenes. arXiv preprint arXiv:2112.08879, 2021.
  29. 29.Chen Sun, Austin Myers, Carl Vondrick, Kevin Murphy, and Cordelia Schmid. Videobert: A joint model for video and language representation learning. In ICCV, pages 7464–7473, 2019.
  30. 30.Antoine Yang, Antoine Miech, Josef Sivic, Ivan Laptev, and Cordelia Schmid. Tubedetr: Spatio-temporal video grounding with transformers. CVPR, 2022.
  31. 31.Marcella Cornia, Matteo Stefanini, Lorenzo Baraldi, and Rita Cucchiara. Meshed-memory transformer for image captioning. In CVPR, pages 10578–10587, 2020.
  32. 32.Luowei Zhou, Hamid Palangi, Lei Zhang, Houdong Hu, Jason Corso, and Jianfeng Gao. Unified vision-language pre-training for image captioning and vqa. In AAAI, volume 34, pages 13041–13049, 2020.
  33. 33.Shizhe Chen, Pierre-Louis Guhur, Cordelia Schmid, and Ivan Laptev. History aware multimodal transformer for vision-and-language navigation. NeurIPS, 34, 2021.
  34. 34.Shizhe Chen, Pierre-Louis Guhur, Makarand Tapaswi, Cordelia Schmid, and Ivan Laptev. Think global, act local: Dual-scale graph transformer for vision-and-language navigation. CVPR, 2022.
  35. 35.Kan Wu, Houwen Peng, Minghao Chen, Jianlong Fu, and Hongyang Chao. Rethinking and improving relative position encoding for vision transformer. In ICCV, pages 10033–10041, 2021.
  36. 36.Geoffrey Hinton, Oriol Vinyals, Jeff Dean, et al. Distilling the knowledge in a neural network. arXiv preprint arXiv:1503.02531, 2(7), 2015.
  37. 37.Lucas Beyer, Xiaohua Zhai, Amélie Royer, Larisa Markeeva, Rohan Anil, and Alexander Kolesnikov. Knowledge distillation: A good teacher is patient and consistent. arXiv preprint arXiv:2106.05237, 2021.
  38. 38.Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu. Tinybert: Distilling bert for natural language understanding. In EMNLP Findings, pages 4163–4174, 2020.
  39. 39.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. NAACL, 2019.
  40. 40.Charles Ruizhongtai Qi, Li Yi, Hao Su, and Leonidas J Guibas. Pointnet++: Deep hierarchical feature learning on point sets in a metric space. NeurIPS, 30, 2017.
  41. 41.Jeffrey Pennington, Richard Socher, and Christopher D Manning. Glove: Global vectors for word representation. In Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP), pages 1532–1543, 2014.
  42. 42.Ilya Loshchilov and Frank Hutter. Fixing weight decay regularization in adam. 2018.
  43. 43.Li Jiang, Hengshuang Zhao, Shaoshuai Shi, Shu Liu, Chi-Wing Fu, and Jiaya Jia. Pointgroup: Dual-set point grouping for 3d instance segmentation. In CVPR, pages 4867–4876, 2020.
  44. 44.Ze Liu, Zheng Zhang, Yue Cao, Han Hu, and Xin Tong. Group-free 3d object detection via transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 2949–2958, 2021.
  45. 45.Charles R Qi, Or Litany, Kaiming He, and Leonidas J Guibas. Deep hough voting for 3d object detection in point clouds. In ICCV, pages 9277–9286, 2019.

Citation

MLA
Chen, S., et al. “Language Conditioned Spatial Relation Reasoning for 3D Object Grounding”. Advances in Neural Information Processing Systems, vol. 35, 2022, pp. 20522–35, https://proceedings.neurips.cc/paper_files/paper/2022/file/819aaee144cb40e887a4aa9e781b1547-Paper-Conference.pdf.
APA
Chen, S., Guhur, P.-L., Tapaswi, M., Schmid, C., & Laptev, I. (2022). Language Conditioned Spatial Relation Reasoning for 3D Object Grounding. Advances in Neural Information Processing Systems, 35, 20522–20535. https://proceedings.neurips.cc/paper_files/paper/2022/file/819aaee144cb40e887a4aa9e781b1547-Paper-Conference.pdf
Chicago
Chen, S., P.-L. Guhur, M. Tapaswi, C. Schmid, and I. Laptev. 2022. “Language Conditioned Spatial Relation Reasoning for 3D Object Grounding”. Advances in Neural Information Processing Systems 35: 20522–35. https://proceedings.neurips.cc/paper_files/paper/2022/file/819aaee144cb40e887a4aa9e781b1547-Paper-Conference.pdf.
Harvard
Chen, S. et al. (2022) “Language Conditioned Spatial Relation Reasoning for 3D Object Grounding”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 20522–20535. Available at: https://proceedings.neurips.cc/paper_files/paper/2022/file/819aaee144cb40e887a4aa9e781b1547-Paper-Conference.pdf.
Vancouver
1. Chen S, Guhur P-L, Tapaswi M, Schmid C, Laptev I (2022) Language Conditioned Spatial Relation Reasoning for 3D Object Grounding. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 20522–20535

BibTeX

@inproceedings{chen2022language,
  title = {Language Conditioned Spatial Relation Reasoning for 3D Object Grounding},
  author = {Chen, Shizhe and Guhur, Pierre-Louis and Tapaswi, Makarand and Schmid, Cordelia and Laptev, Ivan},
  year = {2022},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {35},
  pages = {20522-20535},
  url = {https://proceedings.neurips.cc/paper_files/paper/2022/file/819aaee144cb40e887a4aa9e781b1547-Paper-Conference.pdf}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors