Text2Loc: 3D Point Cloud Localization from Natural Language

Yan XiaLetian ShiZifeng DingJoão F. HenriquesDaniel Cremers

article2024CVPR79 citations

Proposes Text2Loc, a coarse-to-fine framework that localizes natural language descriptions within city-scale 3D point clouds by combining a hierarchical transformer for cross-sentence context with a matching-free fine localization network.

Listen

Autonomous navigation and last-mile delivery services frequently encounter positioning failures in dense urban environments where satellite signals are obstructed by tall buildings and heavy vegetation. In such situations, humans naturally navigate by communicating relative spatial descriptions, such as referencing nearby landmarks or street layouts. Developing systems that can interpret spoken or written descriptions and accurately identify physical positions within pre-mapped 3D environments is essential for advancing autonomous ground vehicles and robotic delivery. The article evaluates a new framework, named Text2Loc, designed to estimate precise geographic coordinates from natural language queries and 3D spatial maps without relying on conventional satellite positioning.

To accomplish this, the system operates through a coarse-to-fine structure. In the coarse global recognition stage, it extracts language features using a pre-trained language model combined with a hierarchical transformer to analyze sentence contexts, while a 3D submap branch extracts spatial, color, and object point density cues. These representations are aligned using contrastive learning to ensure matched text-map pairs cluster closely. In the fine-tuning stage, the system discards the traditional, complex text-to-object matching modules in favor of a matching-free regressor that uses cross-attention transformers and a submap cloning technique. The authors validated the framework on the public KITTI360Pose benchmark, encompassing over 43,000 position-query pairs across 15.5 square kilometers of urban territory.

The experimental findings show that the proposed framework substantially outperforms current leading methods. First, in global place recognition, top-1 submap retrieval recall improved to 0.32 on validation data, exceeding the prior state-of-the-art method by 78%. Second, in fine localization with an error margin under five meters, the system achieved a top-1 recall rate of 0.33 on the test set, outperforming the leading baseline by approximately two times. Third, removing the traditional matching module reduced model parameter count by roughly half and slashed inference runtime from 43.11 milliseconds to 2.27 milliseconds, using only about 5% of the baseline's execution time. Fourth, ablation experiments confirmed that contrastive learning and submap cloning were the primary drivers of retrieval and regression accuracy gains.

These results demonstrate that city-scale natural language localization can be achieved with high accuracy and low computational overhead. Eliminating heavy matching modules significantly lowers computational latency, making deployment on resource-constrained vehicle hardware and edge devices feasible for real-time operations. However, the evaluation indicates that overall accuracy remains heavily dependent on successful coarse place recognition; if the initial step retrieves unrelated map segments due to repetitive urban layouts, the fine regression stage fails. Furthermore, the model exhibits sensitivity to slight phrasing variations in query text. Organizations considering practical adoption should pursue pilot testing on robotic platforms and focus future development on enhancing system robustness against natural speech variations and ambiguous, visually repetitive urban settings.

arXiv: 2311.15977
Cover for Text2Loc: 3D Point Cloud Localization from Natural Language

Abstract

We tackle the problem of 3D point cloud localization based on a few natural linguistic descriptions and introduce a novel neural network, Text2Loc, that fully interprets the semantic relationship between points and text. Text2Loc follows a coarse-to-fine localization pipeline: text-submap global place recognition, followed by fine localization. In global place recognition, relational dynamics among each textual hint are captured in a hierarchical transformer with max-pooling (HTM), whereas a balance between positive and negative pairs is maintained using text-submap contrastive learning. Moreover, we propose a novel matching-free fine localization method to further refine the location predictions, which completely removes the need for complicated text-instance matching and is lighter, faster, and more accurate than previous methods. Extensive experiments show that Text2Loc improves the localization accuracy by up to 2× over the state-of-the-art on the KITTI360Pose dataset. Our project page is publicly available at https://yan-xia.github.io/projects/text2loc/.

Table of Contents

  • 1. Introduction
  • 2. Related work
  • 3. Problem statement
  • 4. Methodology
  • 4.1. Global place recognition
  • 4.2. Fine localization
  • 4.3. Loss function
  • 5. Experiments
  • 5.1. Benchmark Dataset
  • 5.2. Evaluation criteria
  • 5.3. Results
  • 5.3.1 Global place recognition
  • 5.3.2 Fine localization
  • 6. Performance analysis
  • 6.1. Ablation study
  • 6.2. Qualitative analysis
  • 6.3. Computational cost analysis
  • 6.4. Robustness analysis
  • 6.5. Embedding space analysis
  • 7. Conclusion
  • References

Knowls

  1. Knowl 1 — Text2Loc Coarse-to-Fine 3D Point Cloud Localization Formulation

    model/method

    The task of localizing a 3D position from natural language queries in a large-scale 3D environment is formulated as a coarse-to-fine cross-modal mapping problem. The environment is represented by a 3D reference map Mref={mi}i=1M\mathcal{M}_{\text{ref}} = \{m_i\}_{i=1}^M partitioned into cubic submaps mim_i, where each submap contains a set of 3D object instances {Pi,j}j=1p\{P_{i,j}\}_{j=1}^p. The query consists of a text description T={h⃗k}k=1hT = \{\vec{h}_k\}_{k=1}^h specifying spatial relation hints between the target position and surrounding object instances.

    The localization is formulated as minimizing the expected coordinate regression error:

    min⁡ϕ,FE(x,y,T)∼D∥(x,y)−ϕ(T,  arg⁡min⁡m∈Mrefd(F(T),F(m)))∥2\min_{\phi, F} \mathbb{E}_{(x,y,T) \sim \mathcal{D}} \left\| (x, y) - \phi\left(T, \; \arg\min_{m \in \mathcal{M}_{\text{ref}}} d\big(F(T), F(m)\big)\right) \right\|^2

    where:

    • D\mathcal{D} denotes the dataset of query-coordinate tuples ((x,y),T)((x,y), T) with ground-truth 2D planar coordinates (x,y)∈R2(x, y) \in \mathbb{R}^2 in the scene coordinate frame.
    • FF is a dual-branch global embedding function mapping both a textual query TT and a point cloud submap mm into a shared descriptor space R1×C\mathbb{R}^{1 \times C}.
    • d(⋅,⋅)d(\cdot, \cdot) is a distance metric (e.g., Euclidean distance) in the shared descriptor space used to retrieve the nearest candidate submaps.
    • ϕ\phi is a matching-free fine-localization neural network that directly regresses the 2D planar position from the query text representation TT and the retrieved submap mm.
  2. Knowl 2 — Dual-Branch Text and Point Cloud Architecture for Global Place Recognition

    model/method

    Global place recognition embeds textual queries and 3D submaps into a shared metric space R1×C\mathbb{R}^{1 \times C} via two parallel branches:

    1. Text Branch: Textual descriptions are initially passed through a frozen, pre-trained T5 language model. The resulting tokens are fed to a Hierarchical Transformer with Max-pooling (HTM). The HTM consists of residual transformer blocks with Multi-Head Self-Attention (MHSA) and two-layer FeedForward Networks (FFN) with ReLU activations to encode both intra-sentence dependencies and shared inter-sentence contextual relationships, outputting a global text descriptor FT∈R1×CF^T \in \mathbb{R}^{1 \times C}.

    2. 3D Submap Branch: Each object instance PiP_i within a submap SNS_N contains spatial (XYZ) and color (RGB) 6D point coordinates. Semantic features are extracted with PointNet++. In parallel, three separate 3-layer Multi-Layer Perceptrons (MLPs) encode:

      • Instance color from mean RGB values,
      • Instance position from the center coordinates Pˉi\bar{P}_i,
      • Instance point count (number encoder), which supplies category-specific prior information based on object scale.

    The semantic, color, positional, and number embeddings are concatenated and passed through a 3-layer MLP projection layer to produce an instance feature FpiF_{p_i}. All instance descriptors {Fpi}i=1Np\{F_{p_i}\}_{i=1}^{N_p} within the submap are aggregated into a single global submap descriptor FS∈R1×CF^S \in \mathbb{R}^{1 \times C} using a self-attention layer followed by max-pooling.

  3. Knowl 3 — Cross-Modal Contrastive Learning Objective for Global Place Recognition

    equation

    To train the dual-branch encoders for global place recognition while handling the class imbalance between positive and negative submap candidates, Text2Loc employs a symmetric cross-modal InfoNCE contrastive loss over mini-batches.

    For a mini-batch of NN matched text-submap descriptor pairs {(FiT,FiS)}i=1N\{(F^T_i, F^S_i)\}_{i=1}^N with FiT,FiS∈R1×CF^T_i, F^S_i \in \mathbb{R}^{1 \times C}:

    l(i,T,S)=−log⁡exp⁡(FiT⋅FiS/τ)∑j=1Nexp⁡(FiT⋅FjS/τ)−log⁡exp⁡(FiS⋅FiT/τ)∑j=1Nexp⁡(FiS⋅FjT/τ)l(i, T, S) = -\log \frac{\exp(F^T_i \cdot F^S_i / \tau)}{\sum_{j=1}^N \exp(F^T_i \cdot F^S_j / \tau)} - \log \frac{\exp(F^S_i \cdot F^T_i / \tau)}{\sum_{j=1}^N \exp(F^S_i \cdot F^T_j / \tau)}

    L(T,S)=1N∑i=1Nl(i,T,S)L(T, S) = \frac{1}{N} \sum_{i=1}^N l(i, T, S)

    where τ\tau is a temperature hyperparameter and ⋅\cdot denotes the inner product between normalized embeddings.

  4. Knowl 4 — Matching-Free Fine Localization with Cascaded Cross-Attention Transformers

    model/method

    Fine localization refines the position prediction within retrieved submaps without relying on text-to-instance correspondence matching or optimal transport (Sinkhorn) algorithms.

    1. Feature Extraction: Query text is encoded using a frozen pre-trained T5 model followed by an attention unit and max-pooling. Submap point clouds are processed by the instance encoder to produce 3D instance embeddings.
    2. Cascaded Cross-Attention Transformers (CCAT): Two Cascaded Cross-Attention Transformers are stacked sequentially to fuse multi-modal representations:
      • CAT1\text{CAT}_1 takes point cloud features as Query (QQ) and text features as Key (KK) and Value (VV), yielding text-informed point cloud feature maps.
      • CAT2\text{CAT}_2 takes text features as Query (QQ) and the enhanced point features from CAT1\text{CAT}_1 as Key (KK) and Value (VV), generating point-enhanced text features.
    3. Coordinate Regression: Two sequential CCAT blocks fuse the representations, and the final embedding is mapped directly to planar coordinates Cpred=(x^,y^)C_{\text{pred}} = (\hat{x}, \hat{y}) via an MLP.
    4. Loss Function: The fine localization regressor is trained independently using Euclidean mean squared error:

    L(Cgt,Cpred)=∥Cgt−Cpred∥2L(C_{\text{gt}}, C_{\text{pred}}) = \|C_{\text{gt}} - C_{\text{pred}}\|_2

    where Cgt=(x,y)C_{\text{gt}} = (x, y) represents ground-truth 2D coordinates.

  5. Knowl 5 — Prototype-Based Map Cloning (PMC) for Submap Augmentation

    model/method

    Prototype-based Map Cloning (PMC) is a data augmentation method used during fine-localization training to enrich the spatial diversity of candidate submaps around the target location.

    For a training pair consisting of text query TiT_i and candidate submap SiS_i, the collection Gi\mathcal{G}_i of neighboring submap variants is defined as:

    Gi={Sj  |  ∥sˉj−sˉi∥∞<α,  ∥sˉj−ci∥∞<β}\mathcal{G}_i = \left\{ S_j \;\middle|\; \|\bar{s}_j - \bar{s}_i\|_\infty < \alpha, \; \|\bar{s}_j - c_i\|_\infty < \beta \right\}

    where sˉi,sˉj∈R2\bar{s}_i, \bar{s}_j \in \mathbb{R}^2 are the planar centers of submaps SiS_i and SjS_j, ci∈R2c_i \in \mathbb{R}^2 is the ground-truth target coordinate described by TiT_i, α=15 m\alpha = 15\text{ m}, and β=12 m\beta = 12\text{ m}.

    To prevent sampling submaps missing the objects referenced in text TiT_i, an instance mismatch filter is applied with threshold Nm=1N_m = 1, permitting at most one mismatched instance between the textual hints and the submap contents. One submap is uniformly sampled at random from the filtered set Gi\mathcal{G}_i for each training step.

  6. Knowl 6 — KITTI360Pose Localization Recall Benchmark Results

    data/table

    The table below compares 3D localization recall on the KITTI360Pose benchmark across top-kk retrieved submaps (k∈{1,5,10}k \in \{1, 5, 10\}) under error thresholds ϵ<5m\epsilon < 5\text{m}, ϵ<10m\epsilon < 10\text{m}, and ϵ<15m\epsilon < 15\text{m}. KITTI360Pose contains 9 districts covering 43,381 position-query pairs across 15.51 km215.51\text{ km}^2, divided into 5 training scenes (11.59 km211.59\text{ km}^2), 1 validation scene, and 3 testing scenes (2.14 km22.14\text{ km}^2) with 30m×30m30\text{m} \times 30\text{m} cubic submaps at a 10m10\text{m} stride.

    Methods Validation Set Test Set
    k=1k = 1 k=5k = 5 k=10k = 10 k=1k = 1 k=5k = 5 k=10k = 10
    Text2Pos 0.14/0.25/0.31 0.36/0.55/0.61 0.48/0.68/0.74 0.13/0.21/0.25 0.33/0.48/0.52 0.43/0.61/0.65
    RET 0.19/0.30/0.37 0.44/0.62/0.67 0.52/0.72/0.78 0.16/0.25/0.29 0.35/0.51/0.56 0.46/0.65/0.71
    Text2Loc (Ours) 0.37/0.57/0.63 0.68/0.85/0.87 0.77/0.91/0.93 0.33/0.48/0.52 0.61/0.75/0.78 0.71/0.84/0.86

    Text2Loc achieves a top-1 recall of 0.330.33 on the test set for ϵ<5m\epsilon < 5\text{m}, doubling the performance of RET (0.160.16) and outperforming Text2Pos (0.130.13).

  7. Knowl 7 — Submap Retrieval Recall in Global Place Recognition

    data/table

    The table below reports Retrieve Recall at top-kk (k∈{1,3,5}k \in \{1, 3, 5\}) on the KITTI360Pose dataset for text-submap global place recognition. A submap retrieval is counted as correct if the ground-truth target location lies within the retrieved submap bounds.

    Methods Validation Set Test Set
    k=1k = 1 k=3k = 3 k=5k = 5 k=1k = 1 k=3k = 3 k=5k = 5
    Text2Pos 0.14 0.28 0.37 0.12 0.25 0.33
    RET 0.18 0.34 0.44 - - -
    Text2Loc (Ours) 0.32 0.56 0.67 0.28 0.49 0.58

    Text2Loc achieves top-1 retrieval recalls of 0.320.32 on validation and 0.280.28 on test, representing a 78% relative improvement on the validation set over RET (0.180.18).

  8. Knowl 8 — Ablation Study of Global Place Recognition Modules

    data/table

    Ablation experiments on the KITTI360Pose benchmark evaluate the impact of individual components in Text2Loc's global place recognition module:

    • w/o T5: Pre-trained frozen T5 language model is replaced with the Bi-LSTM encoder from Text2Pos.
    • w/o HTM: Hierarchical Transformer with Max-pooling is removed.
    • w/o CL: Symmetric cross-modal contrastive learning is replaced with pairwise ranking loss.
    • w/o NE: Point quantity/number encoder is removed from the 3D submap instance encoder.
    Methods Validation Set Test Set
    k=1k = 1 k=3k = 3 k=5k = 5 k=1k = 1 k=3k = 3 k=5k = 5
    w/o T5 0.29 0.53 0.65 0.26 0.45 0.54
    w/o HTM 0.30 0.54 0.65 0.28 0.48 0.57
    w/o CL 0.21 0.42 0.53 0.20 0.36 0.45
    w/o NE 0.30 0.52 0.63 0.27 0.47 0.56
    Full (Ours) 0.32 0.56 0.67 0.28 0.49 0.58

    Replacing contrastive learning with pairwise ranking loss causes the largest degradation, dropping top-1 validation recall from 0.320.32 to 0.210.21 (a 52%52\% relative drop) and test recall from 0.280.28 to 0.200.20 (a 40%40\% relative drop).

  9. Knowl 9 — Computational Complexity and Ablation of Fine Localization

    data/table

    Evaluation of fine localization components and computational efficiency on the KITTI360Pose dataset, benchmarked on a single NVIDIA TITAN X (12GB) GPU:

    1. Ablation of Fine Localization Modules (Localization Recall ϵ<5m\epsilon < 5\text{m} using identical retrieved coarse submaps):
    Methods Validation Set Test Set
    k=1k = 1 k=5k = 5 k=10k = 10 k=1k = 1 k=5k = 5 k=10k = 10
    Text2Pos (Original) 0.14 0.36 0.48 0.13 0.33 0.43
    Text2Pos* 0.33 0.65 0.75 0.30 0.58 0.67
    Text2Loc_CCAT (w/o PMC) 0.32 0.64 0.74 0.32 0.60 0.70
    Text2Loc_PMC (with Matcher) 0.32 0.64 0.74 0.29 0.56 0.66
    Text2Loc (Full) 0.37 0.68 0.77 0.33 0.61 0.71

    Note: Text2Pos* denotes Text2Pos fine localization evaluated on the submaps retrieved by Text2Loc's coarse stage.

    1. Computational Cost Comparison:
    Methods Parameters (M) Runtime (ms) Localization Recall (k=1,ϵ<5mk=1, \epsilon<5\text{m})
    Text2Loc_Matcher 2.08 43.11 0.30
    Text2Loc (Ours) 1.06 2.27 0.33

    Replacing the SuperGlue/Sinkhorn text-instance matcher with the matching-free CCAT module reduces parameter count from 2.08 M2.08\text{ M} to 1.06 M1.06\text{ M} and runtime from 43.11 ms43.11\text{ ms} to 2.27 ms2.27\text{ ms} (a 19x speedup) while improving top-1 localization recall.

  10. Knowl 10 — Sensitivity to Perturbations in Text Descriptions

    data/table

    Modifying a single descriptive sentence within the input query text significantly degrades localization performance, demonstrating the model's sensitivity to fine-grained linguistic cues.

    Methods Submap Retrieval Recall Localization Recall (k=5k = 5)
    k=1k = 1 k=3k = 3 k=5k = 5 ϵ<5m\epsilon < 5\text{m} ϵ<10m\epsilon < 10\text{m} ϵ<15m\epsilon < 15\text{m}
    Text2Loc_modified 0.15 0.30 0.38 0.39 0.54 0.58
    Text2Loc (Ours) 0.28 0.49 0.58 0.53 0.68 0.71

    When a single sentence in the multi-hint description is altered, top-1 submap retrieval recall on the KITTI360Pose test set drops from 0.280.28 to 0.150.15, and top-5 localization recall at ϵ<5m\epsilon < 5\text{m} decreases from 0.530.53 to 0.390.39.

Coverage note — No substantial contributed material was omitted. Qualitative visual examples and t-SNE figures from the paper were omitted in favor of the complete quantitative tables, algorithmic modules, equations, and ablation studies.

References

  1. 1.Panos Achlioptas, Ahmed Abdelreheem, Fei Xia, Mohamed Elhoseiny, and Leonidas Guibas. Referit3d: Neural listeners for fine-grained 3d object identification in real-world scenes. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part I 16, pages 422–440. Springer, 2020.
  2. 2.Mikaela Angelina Uy and Gim Hee Lee. Pointnetvlad: Deep point cloud based retrieval for large-scale place recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 4470–4479, 2018.
  3. 3.Tiago Barros, Luıs Garrote, Ricardo Pereira, Cristiano Premebida, and Urbano J Nunes. Attdlnet: Attention-based deep network for 3d lidar place recognition. In Iberian Robotics conference, pages 309–320. Springer, 2022.
  4. 4.Dave Zhenyu Chen, Angel X Chang, and Matthias Nießner. Scanrefer: 3d object localization in rgb-d scans using natural language. In European conference on computer vision, pages 202–221. Springer, 2020.
  5. 5.Marco Cuturi. Sinkhorn distances: Lightspeed computation of optimal transport. In Advances in Neural Information Processing Systems. Curran Associates, Inc., 2013.
  6. 6.Haowen Deng, Tolga Birdal, and Slobodan Ilic. Ppfnet: Global context aware local features for robust 3d point matching. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 195–205, 2018.
  7. 7.Gil Elbaz, Tamar Avraham, and Anath Fischer. 3d point cloud registration for localization using a deep neural network auto-encoder. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 4631–4640, 2017.
  8. 8.Zhaoxin Fan, Zhenbo Song, Hongyan Liu, Zhiwu Lu, Jun He, and Xiaoyong Du. Svt-net: Super light-weight sparse voxel transformer for large scale place recognition. AAAI, 2022.
  9. 9.Mingtao Feng, Zhen Li, Qi Li, Liang Zhang, XiangDong Zhang, Guangming Zhu, Hui Zhang, Yaonan Wang, and Ajmal Mian. Free-form description guided 3d visual graph network for object grounding in point cloud. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 3722–3731, 2021.
  10. 10.Sepp Hochreiter and Jürgen Schmidhuber. Long short-term memory. Neural computation, 9(8):1735–1780, 1997.
  11. 11.Yihan Hu, Jiazhi Yang, Li Chen, Keyu Li, Chonghao Sima, Xizhou Zhu, Siqi Chai, Senyao Du, Tianwei Lin, Wenhai Wang, Lewei Lu, Xiaosong Jia, Qiang Liu, Jifeng Dai, Yu Qiao, and Hongyang Li. Planning-oriented autonomous driving. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023.
  12. 12.Manuel Kolmet, Qunjie Zhou, Aljoša Ošep, and Laura Leal-Taixé. Text2pos: Text-to-point-cloud cross-modal localization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6687–6696, 2022.
  13. 13.Jacek Komorowski. Minkloc3d: Point cloud based large-scale place recognition. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 1790–1799, 2021.
  14. 14.Tsung-Yi Lin, Piotr Dollár, Ross Girshick, Kaiming He, Bharath Hariharan, and Serge Belongie. Feature pyramid networks for object detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 2117–2125, 2017.
  15. 15.David G Lowe. Distinctive image features from scale-invariant keypoints. International journal of computer vision, 60:91–110, 2004.
  16. 16.Junyi Ma, Jun Zhang, Jintao Xu, Rui Ai, Weihao Gu, and Xieyuanli Chen. Overlaptransformer: An efficient and yaw-angle-invariant transformer network for lidar-based place recognition. IEEE Robotics and Automation Letters, 7(3): 6958–6965, 2022.
  17. 17.Junyi Ma, Guangming Xiong, Jingyi Xu, and Xieyuanli Chen. Cvtnet: A cross-view transformer network for place recognition using lidar data. arXiv preprint arXiv:2302.01665, 2023.
  18. 18.Zhixiang Min, Bingbing Zhuang, Samuel Schulter, Buyu Liu, Enrique Dunn, and Manmohan Chandraker. Neurocs: Neural nocs supervision for monocular 3d object localization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 21404–21414, 2023.
  19. 19.Mihir Prabhudesai, Hsiao-Yu Fish Tung, Syed Ashar Javed, Maximilian Sieb, Adam W Harley, and Katerina Fragkiadaki. Embodied language grounding with implicit 3d visual feature representations. 2019.
  20. 20.Charles Ruizhongtai Qi, Li Yi, Hao Su, and Leonidas J Guibas. Pointnet++: Deep hierarchical feature learning on point sets in a metric space. Advances in neural information processing systems, 30, 2017.
  21. 21.Filip Radenović, Giorgos Tolias, and Ondřej Chum. Fine-tuning cnn image retrieval with no human annotation. IEEE Transactions on Pattern Analysis and Machine Intelligence, 41(7):1655–1668, 2018.
  22. 22.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International conference on machine learning, pages 8748–8763. PMLR, 2021.
  23. 23.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21(140):1–67, 2020.
  24. 24.Ethan Rublee, Vincent Rabaud, Kurt Konolige, and Gary Bradski. Orb: An efficient alternative to sift or surf. In 2011 International conference on computer vision, pages 2564–2571. Ieee, 2011.
  25. 25.Haşim Sak, Andrew W. Senior, and Françoise Beaufays. Long short-term memory based recurrent neural network architectures for large vocabulary speech recognition. CoRR, abs/1402.1128, 2014.
  26. 26.Paul-Edouard Sarlin, Cesar Cadena, Roland Siegwart, and Marcin Dymczyk. From coarse to fine: Robust hierarchical localization at large scale. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12716–12725, 2019.
  27. 27.Paul-Edouard Sarlin, Daniel DeTone, Tomasz Malisiewicz, and Andrew Rabinovich. SuperGlue: Learning feature matching with graph neural networks. In CVPR, 2020.
  28. 28.Paul-Edouard Sarlin, Daniel DeTone, Tsun-Yi Yang, Armen Avetisyan, Julian Straub, Tomasz Malisiewicz, Samuel Rota Bulò, Richard Newcombe, Peter Kontschieder, and Vasileios Balntas. Orienternet: Visual localization in 2d public maps with neural matching. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 21632–21642, 2023.
  29. 29.Torsten Sattler, Bastian Leibe, and Leif Kobbelt. Efficient & effective prioritized matching for large-scale image-based localization. IEEE transactions on pattern analysis and machine intelligence, 39(9):1744–1756, 2016.
  30. 30.Aleksandr Segal, Dirk Haehnel, and Sebastian Thrun. Generalized-icp. In Robotics: science and systems, page 435. Seattle, WA, 2009.
  31. 31.Laurens Van der Maaten and Geoffrey Hinton. Visualizing data using t-sne. JMLR, 2008.
  32. 32.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in neural information processing systems, 30, 2017.
  33. 33.Guangzhi Wang, Hehe Fan, and Mohan Kankanhalli. Text to point cloud localization with relation-enhanced transformer. arXiv preprint arXiv:2301.05372, 2023.
  34. 34.Yan Xia. Perception of vehicles and place recognition in urban environment based on MLS point clouds. PhD thesis, Technische Universität München, 2023.
  35. 35.Yan Xia, Yusheng Xu, Shuang Li, Rui Wang, Juan Du, Daniel Cremers, and Uwe Stilla. Soe-net: A self-attention and orientation encoding network for point cloud based place recognition. In Proceedings of the IEEE/CVF Conference on computer vision and pattern recognition, pages 11348–11357, 2021.
  36. 36.Yan Xia, Yusheng Xu, Cheng Wang, and Uwe Stilla. Vpc-net: Completion of 3d vehicles from mls point clouds. ISPRS Journal of Photogrammetry and Remote Sensing, 174:166–181, 2021.
  37. 37.Yan Xia, Mariia Gladkova, Rui Wang, Qianyun Li, Uwe Stilla, João F Henriques, and Daniel Cremers. Casspr: Cross attention single scan place recognition. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 8461–8472, 2023.
  38. 38.Yan Xia, Qiangqiang Wu, Wei Li, Antoni B Chan, and Uwe Stilla. A lightweight and detector-free 3d single object tracker on point clouds. IEEE Transactions on Intelligent Transportation Systems, 2023.
  39. 39.Zhihao Yuan, Xu Yan, Yinghong Liao, Ruimao Zhang, Sheng Wang, Zhen Li, and Shuguang Cui. Instancerefer: Cooperative holistic understanding for visual grounding on point clouds through instance multi-level contextual referring. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 1791–1800, 2021.
  40. 40.Wenxiao Zhang, Huajian Zhou, Zhen Dong, Qingan Yan, and Chunxia Xiao. Rank-pointretrieval: Reranking point cloud retrieval via a visually consistent registration evaluation. IEEE Transactions on Visualization and Computer Graphics, 2022.
  41. 41.Zhicheng Zhou, Cheng Zhao, Daniel Adolfsson, Songzhi Su, Yang Gao, Tom Duckett, and Li Sun. Ndt-transformer: Large-scale 3d point cloud localisation using the normal distribution transform representation. In 2021 IEEE International Conference on Robotics and Automation (ICRA), pages 5654–5660. IEEE, 2021.

Citation

MLA
Xia, Y., et al. “Text2Loc: 3D Point Cloud Localization from Natural Language”. arXiv, 2023, http://arxiv.org/abs/2311.15977v2.
APA
Xia, Y., Shi, L., Ding, Z., Henriques, J. F., & Cremers, D. (2023). Text2Loc: 3D Point Cloud Localization from Natural Language. arXiv. http://arxiv.org/abs/2311.15977v2
Chicago
Xia, Y., L. Shi, Z. Ding, J. F. Henriques, and D. Cremers. 2023. “Text2Loc: 3D Point Cloud Localization from Natural Language”. arXiv. http://arxiv.org/abs/2311.15977v2.
Harvard
Xia, Y. et al. (2023) “Text2Loc: 3D Point Cloud Localization from Natural Language”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2311.15977v2.
Vancouver
1. Xia Y, Shi L, Ding Z, Henriques JF, Cremers D (2023) Text2Loc: 3D Point Cloud Localization from Natural Language. arXiv

BibTeX

@article{xia2023text2loc,
  title = {Text2Loc: 3D Point Cloud Localization from Natural Language},
  author = {Xia, Yan and Shi, Letian and Ding, Zifeng and Henriques, João F. and Cremers, Daniel},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2311.15977v2},
  eprint = {2311.15977}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE