Built independently by an author, for readers. Read the story and support ChapterPal

keyword

text-instance matching

Text-instance matching is a multimodal alignment process in computer vision and natural language processing that establishes direct correspondences between linguistic elements in a text description and specific object instances identified within visual or spatial data. Commonly employed in tasks such as visual grounding and cross-modal spatial localization, this technique maps individual words or referring expressions to discrete segmented objects, bounding boxes, or spatial landmarks in a scene representation, such as a 3D point cloud or image. By resolving these fine-grained entity-level relationships, systems can verify spatial contexts, determine object-level associations, and refine coordinate or position estimates based on detailed descriptive cues.

1 item

Text2Loc: 3D Point Cloud Localization from Natural Language

Text2Loc: 3D Point Cloud Localization from Natural Language

Yan Xia, Letian Shi, Zifeng Ding, João F. Henriques, Daniel Cremers

OrganizationsLudwig Maximilian University of MunichMunich Center for Machine LearningTechnical University of MunichUniversity of Oxford

Why you should read this

Proposes Text2Loc, a coarse-to-fine framework that localizes natural language descriptions within city-scale 3D point clouds by combining a hierarchical transformer for cross-sentence context with a matching-free fine localization network.

We tackle the problem of 3D point cloud localization based on a few natural linguistic descriptions and introduce a novel neural network, Text2Loc, that fully interprets the semantic relationship between points and text. Text2Loc follows a coarse-to-fine localization pipeline: text-submap global place recognition, followed by fine localization. In global place recognition, relational dynamics among each textual hint are captured in a hierarchical transformer with max-pooling (HTM), whereas a balance between positive and negative pairs is maintained using text-submap contrastive learning. Moreover, we propose a novel matching-free fine localization method to further refine the location predictions, which completely removes the need for complicated text-instance matching and is lighter, faster, and more accurate than previous methods. Extensive experiments show that Text2Loc improves the localization accuracy by up to 2× over the state-of-the-art on the KITTI360Pose dataset. Our project page is publicly available at https://yan-xia.github.io/projects/text2loc/.

Added

2026-09-26