Multi-View Transformer for 3D Visual Grounding

Shijia HuangYilun ChenJiaya JiaLiwei Wang

article2022CVPR208 citations

Proposes a Multi-View Transformer that projects 3D point cloud coordinates across rotated views to aggregate spatial information simultaneously, eliminating viewpoint bias and achieving state-of-the-art performance on 3D visual grounding benchmarks.

Listen

Visual grounding in three dimensions requires automated systems to interpret natural language descriptions and locate the corresponding target objects within complex 3D scenes. This capability is critical for emerging autonomous robotics, intelligent virtual agents, and spatial navigation. A core operational challenge is that spatial language, such as "on the left" or "facing the bed," frequently depends on an implicit human vantage point. When an autonomous system views a scene from an orientation different from the speaker's, traditional models often fail because their spatial position representations are tied to a single, static view.

The article demonstrates and evaluates the Multi-View Transformer, a novel framework designed to overcome vantage-point discrepancies by learning a view-robust multi-modal representation for 3D visual grounding.

The evaluated approach projects a 3D scene into a multi-view space through systematic equal-angle rotations around the vertical axis. To keep computational costs minimal, the architecture decouples processing by extracting point-cloud features once and sharing them across all views, updating only object position coordinates per perspective. Object representations and language queries are fused across views using a transformer decoder, integrated via average pooling, and further refined through an auxiliary language-guided classification task. The framework was evaluated across standard benchmark datasets, including Nr3D, Sr3D, and ScanRefer, using scenes and object descriptions derived from real-world indoor environments.

The experimental findings show substantial performance improvements across all benchmarks. On the Nr3D benchmark, the Multi-View Transformer achieved an overall accuracy of 55.1%, outperforming the best competing method by 11.2 percentage points and surpassing approaches that rely on additional 2D semantic data by 5.9 percentage points. On view-dependent queries within Nr3D, accuracy reached 54.3%, representing an improvement of 12.6 percentage points over the prior best single-view baseline. On the Sr3D benchmark, the framework attained 64.5% accuracy, exceeding previous methods by 7.1 percentage points without requiring 2D supervision. Furthermore, these accuracy gains required negligible computational overhead, increasing per-query inference latency by only 3 milliseconds—from 261 milliseconds under a single view to 264 milliseconds under a four-view configuration.

These results establish that modeling multi-view coordinate spaces directly within neural representations resolves view ambiguity more effectively than simply augmenting training data with random rotations or attempting explicit viewpoint prediction. By rendering systems invariant to starting orientations, this approach eliminates the need for speakers to provide explicit viewing cues, substantially improving reliability for real-world robotic and navigation tasks.

Organizations developing spatial artificial intelligence and embodied robotics should adopt multi-view spatial encoding strategies and incorporate multi-modal classification objectives into their visual grounding pipelines. A four-view configuration provides the optimal balance, capturing necessary spatial context without introducing the computational redundancy or training instability observed at higher view counts.

Confidence in these findings is high across standard indoor benchmark conditions. However, performance boundaries remain constrained in environments featuring highly intricate spatial references, heavy object clutter, or dense distractors where visual recognition errors can still occur. Future development should evaluate deployment in unstructured outdoor environments and test performance under dynamic, real-time viewpoint shifts.

arXiv: 2204.02174
Cover for Multi-View Transformer for 3D Visual Grounding

Abstract

The 3D visual grounding task aims to ground a natural language description to the targeted object in a 3D scene, which is usually represented in 3D point clouds. Previous works studied visual grounding under specific views. The vision-language correspondence learned by this way can easily fail once the view changes. In this paper, we propose a Multi-View Transformer (MVT) for 3D visual grounding. We project the 3D scene to a multi-view space, in which the position information of the 3D scene under different views are modeled simultaneously and aggregated together. The multi-view space enables the network to learn a more robust multi-modal representation for 3D visual grounding and eliminates the dependence on specific views. Extensive experiments show that our approach significantly outperforms all state-of-the-art methods. Specifically, on Nr3D and Sr3D datasets, our method outperforms the best competitor by 11.2% and 7.1% and even surpasses recent work with extra 2D assistance by 5.9% and 6.6%. Our code is available at https://github.com/sega-hsj/MVT-3DVG.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Method
  • 3.1. Multi-View 3D Visual Grounding
  • 3.2. Object Feature Encoding
  • 3.3. Multi-Modal Feature Fusion
  • 3.4. Multi-View Aggregation
  • 3.5. Improving Object Encoder
  • 3.6. Overall Loss Functions
  • 4. Experiments
  • 4.1. Datasets
  • 4.2. Experimental Setting
  • 4.3. 3D Visual Grounding Results
  • 4.4. Ablation studies
  • 4.4.1 Effectiveness of each Component
  • 4.4.2 Effectiveness of Multi-View Modeling.
  • 5. Conclusion
  • References

Knowls

  1. Knowl 1 — View-invariant multi-view scene representation

    model/method

    The Multi-View Transformer (MVT) addresses 3D visual grounding by transforming a scene into several equally spaced views before learning the vision-language correspondence. Let SS be a 3D scene under an arbitrary starting view, QQ be a natural-language query, and NN be the number of views. The jj-th rotated scene is

    Sj=R(θvj)×S,S^{j}=R(\theta_v^{j})\times S,

    where the clockwise rotation about the vertical ZZ axis is

    R(θ)=(cos⁡θ−sin⁡θ0sin⁡θcos⁡θ0001)T,θvj=2π(j−1)N,j∈{1,…,N}.R(\theta)= \begin{pmatrix} \cos\theta & -\sin\theta & 0\\ \sin\theta & \cos\theta & 0\\ 0 & 0 & 1 \end{pmatrix}^{T}, \qquad \theta_v^{j}=\frac{2\pi(j-1)}{N},\quad j\in\{1,\ldots,N\}.

    Here, θ\theta is the rotation angle in radians and TT denotes matrix transpose. The first view is the original scene because θv1=0\theta_v^{1}=0. A shared feature-extraction network ff processes every rotated scene together with QQ, and an order-independent aggregation function gg produces the grounding representation:

    VG(S,Q)=g(f(S1,Q),…,f(SN,Q)).\mathrm{VG}(S,Q)=g\big(f(S^{1},Q),\ldots,f(S^{N},Q)\big).

    Changing the starting view only permutes the same set of rotated views, so the final representation is invariant to the initial view when ff is shared and gg is order-independent. MVT uses N∈{1,2,4,8}N\in\{1,2,4,8\} as candidate view counts and uses N=4N=4 by default.

  2. Knowl 2 — Decoupled object encoding across views

    model/method

    MVT separates view-independent point-cloud appearance features from view-dependent positional features so that expensive point-cloud processing is performed only once. Suppose a scene contains MM candidate objects. For object ii, MVT samples 1024 RGB-XYZ points, represented as pci∈R1024×6pc_i\in\mathbb{R}^{1024\times 6}, and computes a shared point-cloud feature

    xi=LN(Wx PointNet++(pci)),x_i=\mathrm{LN}\big(W_x\,\mathrm{PointNet++}(pc_i)\big),

    where xi∈Rdx_i\in\mathbb{R}^{d} is the appearance feature, WxW_x is a learned projection matrix, and LN\mathrm{LN} is layer normalization. Let bi∈R3b_i\in\mathbb{R}^{3} be the object box center and let rir_i be its scalar box-size descriptor. Under view jj, the center becomes

    bij=R(θvj)×bi,b_i^{j}=R(\theta_v^{j})\times b_i,

    and its dd-dimensional positional encoding is

    PE(bij)=LN(Wb[bij,ri]),\mathrm{PE}(b_i^{j})=\mathrm{LN}\big(W_b[b_i^{j},r_i]\big),

    where WbW_b is a learned projection matrix and [bij,ri][b_i^{j},r_i] denotes concatenation. The final feature for object ii in view jj is

    oij=xi+PE(bij).o_i^{j}=x_i+\mathrm{PE}(b_i^{j}).

    Thus, xix_i is shared across all views and captures visual properties such as category and color, while PE(bij)\mathrm{PE}(b_i^{j}) changes with the view and supplies the corresponding position and size information. This factorization substantially reduces the cost of constructing the multi-view representation.

  3. Knowl 3 — Transformer fusion and late multi-view aggregation

    model/method

    For a query QQ containing k1k_1 words, MVT uses a fine-tuned BERT encoder to produce a sentence feature ls∈Rdl_s\in\mathbb{R}^{d} and word features l1,…,lk1∈Rdl_1,\ldots,l_{k_1}\in\mathbb{R}^{d}:

    {ls,l1,…,lk1}=BERT(Q).\{l_s,l_1,\ldots,l_{k_1}\}=\mathrm{BERT}(Q).

    For each view jj, the object sequence Oj={o1j,…,oMj}O^j=\{o_1^j,\ldots,o_M^j\} is fused with the language sequence L={ls,l1,…,lk1}L=\{l_s,l_1,\ldots,l_{k_1}\} using the same transformer decoder across all views:

    Fj=Decoder(Oj,L),F^j=\mathrm{Decoder}(O^j,L),

    where Fj={f1j,…,fMj}F^j=\{f_1^j,\ldots,f_M^j\} contains one multimodal feature per candidate object. Each decoder layer first applies self-attention to model relationships among objects and then applies cross-attention between object and language features. MVT aggregates the fused features for each object by averaging across the NN views:

    gi=1N∑j=1Nfij,G={g1,…,gM}.g_i=\frac{1}{N}\sum_{j=1}^{N}f_i^j, \qquad G=\{g_1,\ldots,g_M\}.

    Two fully connected layers transform GG into one grounding score per candidate object; the object with the highest score is selected. Aggregating after multimodal fusion preserves view-specific positional information until it has been aligned with language.

  4. Knowl 4 — Language-guided object classification and joint training objective

    model/method

    MVT uses category names as language supervision to improve the object encoder. Let k2k_2 be the number of available object categories and let ct∈Rdc_t\in\mathbb{R}^{d} be the sentence-level BERT feature of category-name text tt. The category feature matrix is C={c1,…,ck2}∈Rk2×dC=\{c_1,\ldots,c_{k_2}\}\in\mathbb{R}^{k_2\times d}. For object ii, whose view-independent point-cloud feature is xi∈Rdx_i\in\mathbb{R}^{d}, category logits are computed by inner products:

    pi=Cxi∈Rk2.p_i=Cx_i\in\mathbb{R}^{k_2}.

    A cross-entropy loss Lobj\mathcal{L}_{obj} supervises these logits using the object's category label. In parallel, a two-layer fully connected classifier applied to the query sentence feature lsl_s predicts the category described by the query, with cross-entropy loss Ltext\mathcal{L}_{text}. The final grounding scores are supervised by a cross-entropy loss Lref\mathcal{L}_{ref} over candidate objects. The complete training objective is

    Ltotal=Lref+α(Ltext+Lobj),\mathcal{L}_{total}=\mathcal{L}_{ref}+\alpha\big(\mathcal{L}_{text}+\mathcal{L}_{obj}\big),

    where MVT uses α=0.5\alpha=0.5. The category vocabulary contains 524 classes for Nr3D and 607 classes for Sr3D. This auxiliary language-guided classification explicitly aligns object features with fine-grained category semantics rather than training the object encoder from visual data alone.

  5. Knowl 5 — Datasets, evaluation protocol, and MVT training configuration

    experimental setup

    MVT is evaluated on three indoor-scene 3D visual-grounding benchmarks. Nr3D contains 45,503 human utterances over 707 ScanNet scenes, 76 fine-grained object classes, and at most six same-class distractors; its reported splits are Easy, Hard, View-dependent, and View-independent. Sr3D contains 83,572 template-generated utterances, while Sr3D+ expands this to 114,532 utterances. ScanRefer contains 51,583 human utterances over 800 ScanNet scenes, divided into 36,665 training, 9,508 validation, and 5,410 test samples, with Unique and Multiple subsets based on whether same-class objects are present.

    For Nr3D, Sr3D, and Sr3D+, proposals are ground-truth objects and performance is classification accuracy. For ScanRefer, proposals come from a pretrained PointGroup detector and performance is Acc@mIoU\mathrm{Acc}@m\mathrm{IoU} for m∈{0.25,0.5}m\in\{0.25,0.5\}, namely the fraction of queries whose predicted 3D box has IoU greater than mm with the ground-truth box.

    The implementation uses feature dimension d=768d=768, the first three layers of BERTBASE as the text encoder, a four-layer transformer decoder trained from scratch, and four views (N=4N=4). Adam optimization uses batch size 24 and initial learning rate 0.00050.0005; the transformer learning rate is multiplied by 0.10.1. After epoch 40, the learning rate is multiplied by 0.650.65 every 10 epochs until 100 epochs.

  6. Knowl 6 — State-of-the-art grounding accuracy on Nr3D and Sr3D

    data/table

    On the ground-truth-proposal benchmarks, MVT substantially improves over methods without extra 2D assistance and also exceeds SAT, which uses 2D semantic assistance during training. The entries below are grounding accuracy percentages; ±\pm values are the variations reported by the paper.

    Dataset/training Method Overall Easy Hard View-dep. View-indep.
    Nr3D LanguageRefer 43.9% 51.0% 36.6% 41.7% 45.0%
    Nr3D SAT (2D assist.) 49.2%±\pm0.3% 56.3%±\pm0.5% 42.4%±\pm0.4% 46.9%±\pm0.3% 50.4%±\pm0.3%
    Nr3D MVT 55.1%±\pm0.3% 61.3%±\pm0.4% 49.1%±\pm0.4% 54.3%±\pm0.5% 55.4%±\pm0.3%
    Nr3D w/ Sr3D TransRefer3D 47.2%±\pm0.3% 55.4%±\pm0.5% 39.3%±\pm0.5% 40.3%±\pm0.4% 50.6%±\pm0.2%
    Nr3D w/ Sr3D SAT (2D assist.) 53.9%±\pm0.2% 61.5%±\pm0.1% 46.7%±\pm0.3% 52.7%±\pm0.7% 54.5%±\pm0.3%
    Nr3D w/ Sr3D MVT 58.5%±\pm0.2% 65.6%±\pm0.2% 51.6%±\pm0.3% 56.6%±\pm0.3% 59.4%±\pm0.2%
    Nr3D w/ Sr3D+ TransRefer3D 48.0%±\pm0.2% 56.7%±\pm0.4% 39.6%±\pm0.2% 42.5%±\pm0.2% 50.7%±\pm0.4%
    Nr3D w/ Sr3D+ SAT (2D assist.) 56.5%±\pm0.1% 64.9%±\pm0.2% 48.4%±\pm0.1% 54.4%±\pm0.3% 57.6%±\pm0.1%
    Nr3D w/ Sr3D+ MVT 59.5%±\pm0.2% 67.4%±\pm0.4% 52.7%±\pm0.4% 59.1%±\pm0.5% 60.3%±\pm0.1%
    Sr3D TransRefer3D 57.4%±\pm0.2% 60.5%±\pm0.2% 50.2%±\pm0.2% 49.9%±\pm0.6% 57.7%±\pm0.2%
    Sr3D SAT (2D assist.) 57.9% 61.2% 50.0% 49.2% 58.3%
    Sr3D MVT 64.5%±\pm0.1% 66.9%±\pm0.1% 58.8%±\pm0.1% 58.4%±\pm0.8% 64.7%±\pm0.1%

    MVT improves over the best competitor without extra assistance by 11.2 percentage points on Nr3D, from 43.9% to 55.1%, and by 7.1 points on Sr3D, from 57.4% to 64.5%. It also exceeds SAT by 5.9 points on Nr3D and 6.6 points on Sr3D. Training with the synthetic Sr3D or Sr3D+ data further raises Nr3D accuracy to 58.5% and 59.5%, respectively.

  7. Knowl 7 — ScanRefer performance and inference cost

    data/table

    On ScanRefer, where object proposals and initial object features are supplied by a pretrained detector, MVT tests the benefit of multi-view representation on top of the detector pipeline. The table reports Acc@0.25\mathrm{Acc}@0.25 and Acc@0.5\mathrm{Acc}@0.5 for Unique, Multiple, and Overall query subsets.

    Method Unique Multiple Overall
    [email protected] [email protected] [email protected] [email protected] [email protected] [email protected]
    ScanRefer 65.00% 43.31% 30.63% 19.75% 37.30% 24.32%
    TGNN 64.50% 53.01% 27.01% 21.88% 34.29% 27.92%
    TGNN + BERT 68.61% 56.80% 29.84% 23.18% 37.37% 29.70%
    InstanceRefer 77.45% 66.83% 31.27% 24.77% 40.23% 32.93%
    MVT, view number = 1 74.36% 63.04% 29.65% 23.44% 38.33% 31.12%
    MVT, view number = 4 77.67% 66.45% 31.92% 25.26% 40.80% 33.26%

    Using four views rather than one raises Overall Acc@0.25\mathrm{Acc}@0.25 from 38.33% to 40.80% and Overall Acc@0.5\mathrm{Acc}@0.5 from 31.12% to 33.26%; it also improves both Unique and Multiple subsets. On a TITAN X Pascal, average inference time is 261 ms per scene-query pair for one view and 264 ms for four views, so the reported multi-view gain adds only 3 ms in this configuration.

  8. Knowl 8 — Component ablation demonstrates complementary MVT gains

    empirical result

    An ablation on Nr3D isolates the transformer decoder, random rotation augmentation, language-guided supervision, and multi-view encoding. Accuracy is reported as a percentage across the overall, Easy, Hard, View-dependent, and View-independent splits. The baseline uses linear object processing followed by direct multiplicative fusion with language and does not model object-object relationships. Rotation augmentation samples one of the four equal-angle rotations at random.

    Configuration Overall Easy Hard View-dep. View-indep.
    Baseline 36.9% 43.5% 30.5% 32.9% 38.8%
    + transformer decoder 40.4% 47.5% 33.7% 38.4% 41.4%
    + rotation augmentation 40.8% 48.4% 33.5% 35.2% 43.6%
    + language-guided supervision 46.2% 53.8% 38.9% 42.6% 48.0%
    + multi-view encoding instead of language-guided supervision 52.3% 59.1% 45.7% 50.3% 53.3%
    + language-guided supervision and multi-view encoding, without rotation augmentation 53.3% 59.8% 47.0% 54.4% 52.7%
    + language-guided supervision, multi-view encoding, and rotation augmentation 55.1% 61.3% 49.1% 54.3% 55.4%

    The decoder raises overall accuracy from 36.9% to 40.4% by modeling object relationships and vision-language fusion. Rotation augmentation alone does not remove view dependence and reduces View-dependent accuracy from 38.4% to 35.2%. Language-guided supervision raises overall accuracy to 46.2%, while multi-view encoding raises it to 52.3% and produces the largest robustness gain on View-dependent queries. Combining all components reaches 55.1% overall accuracy.

  9. Knowl 9 — Four views and late averaging are the most effective multi-view design

    empirical result

    MVT's multi-view design was evaluated on Nr3D by varying the number of views used for training and testing. Accuracy is shown for the overall, Easy, Hard, View-dependent, and View-independent splits.

    Train views Test views Overall Easy Hard View-dep. View-indep.
    4 1 51.6% 58.4% 45.0% 50.7% 52.0%
    4 2 54.3% 60.6% 48.3% 53.7% 54.7%
    4 4 55.1% 61.3% 49.1% 54.3% 55.4%
    4 8 54.8% 61.1% 48.8% 54.2% 55.1%
    1 1 46.2% 53.8% 38.9% 42.6% 48.0%
    2 2 51.9% 60.2% 43.9% 50.7% 52.5%
    8 8 53.4% 60.9% 46.2% 53.5% 53.4%

    Training with four views but testing with one already gives 51.6% overall accuracy, versus 46.2% when both training and testing use one view, showing that multi-view training improves the learned representation even under a single-view test. Four views are sufficient for the best reported matched train/test result; eight views give slightly lower accuracy, which the paper attributes to redundant views and reduced training efficiency.

    The aggregation stage was also varied on Nr3D:

    Aggregation choice Overall Easy Hard View-dep. View-indep.
    After positional encoding 42.5% 49.5% 35.8% 37.4% 45.0%
    After object feature generation 47.4% 54.2% 41.0% 44.3% 49.0%
    After multimodal fusion 55.1% 61.3% 49.1% 54.3% 55.4%
    Average aggregation 55.1% 61.3% 49.1% 54.3% 55.4%
    Max aggregation 52.7% 60.0% 45.7% 51.7% 53.2%
    Average plus max 55.2% 62.1% 48.5% 54.4% 55.6%

    Averaging positional encodings too early destroys useful position information, while averaging after object feature generation retains position but remains inferior to averaging after multimodal fusion. Average aggregation is more effective than max aggregation; average-plus-max produces only a marginal change relative to average alone.

  10. Knowl 10 — Failure modes under complex spatial language

    limitation

    The paper reports that MVT remains vulnerable to complicated language queries combining fine-grained categories with multiple spatial relations. In qualitative failure cases, the model predicts an object of the wrong category or an object at the wrong position, even though multi-view modeling improves the correspondence between scene positions and language. The method therefore reduces, but does not eliminate, errors caused by ambiguous or compositionally complex spatial descriptions.

Coverage note — Qualitative success visualizations and individual example images were omitted as they illustrate the quantitative findings and failure mode without adding a separate standalone contribution.

References

  1. 1.Panos Achlioptas, Ahmed Abdelreheem, Fei Xia, Mohamed Elhoseiny, and Leonidas Guibas. Referit3d: Neural listeners for fine-grained 3d object identification in real-world scenes. In ECCV, 2020. 1, 2, 3, 5, 6, 7
  2. 2.Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton. Layer normalization. arXiv, 2016. 4, 8
  3. 3.Dave Zhenyu Chen, Angel X Chang, and Matthias Nießner. Scanrefer: 3d object localization in rgb-d scans using natural language. In ECCV, 2020. 1, 2, 3, 5, 6, 7
  4. 4.Xinlei Chen, Hao Fang, Tsung-Yi Lin, Ramakrishna Vedantam, Saurabh Gupta, Piotr Dollar, and C Lawrence Zitnick. Microsoft coco captions: Data collection and evaluation server. arXiv, 2015. 2
  5. 5.Xiaozhi Chen, Huimin Ma, Ji Wan, Bo Li, and Tian Xia. Multi-view 3d object detection network for autonomous driving. In CVPR, 2017. 3
  6. 6.Junyoung Chung, Caglar Gulcehre, KyungHyun Cho, and Yoshua Bengio. Empirical evaluation of gated recurrent neural networks on sequence modeling. NIPS, 2014. 2
  7. 7.Angela Dai, Angel X Chang, Manolis Savva, Maciej Halber, Thomas Funkhouser, and Matthias Nießner. Scannet: Richly-annotated 3d reconstructions of indoor scenes. In CVPR, 2017. 2, 5
  8. 8.Karan Desai and Justin Johnson. Virtex: Learning visual representations from textual annotations. In CVPR, 2021. 2
  9. 9.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. In NAACL, 2019. 4, 5, 7
  10. 10.Mingtao Feng, Zhen Li, Qi Li, Liang Zhang, XiangDong Zhang, Guangming Zhu, Hui Zhang, Yaonan Wang, and Ajmal Mian. Free-form description guided 3d visual graph network for object grounding in point cloud. arXiv, 2021. 2, 6
  11. 11.Qi Feng, Vitaly Ablavsky, and Stan Sclaroff. Cityflow-nl: Tracking and retrieval of vehicles at city scale by natural language descriptions. arXiv, 2021. 1
  12. 12.Andreas Geiger, Philip Lenz, and Raquel Urtasun. Are we ready for autonomous driving? the kitti vision benchmark suite. In CVPR, 2012. 3
  13. 13.Lluis Gomez, Yash Patel, Marçal Rusinol, Dimosthenis Karatzas, and CV Jawahar. Self-supervised learning of visual features through embedding images into text topic spaces. In CVPR, 2017. 2
  14. 14.Dailan He, Yusheng Zhao, Junyu Luo, Tianrui Hui, Shaofei Huang, Aixi Zhang, and Si Liu. Transrefer3d: Entity-and-relation aware transformer for fine-grained 3d visual grounding. In Proceedings of the 29th ACM International Conference on Multimedia, 2021. 1, 2, 4, 6
  15. 15.Pin-Hao Huang, Han-Hung Lee, Hwann-Tzong Chen, and Tyng-Luh Liu. Text-guided graph neural networks for referring 3d instance segmentation. In AAAI, 2021. 7
  16. 16.Li Jiang, Hengshuang Zhao, Shaoshuai Shi, Shu Liu, Chi-Wing Fu, and Jiaya Jia. Pointgroup: Dual-set point grouping for 3d instance segmentation. In CVPR, 2020. 2, 5
  17. 17.Armand Joulin, Laurens Van Der Maaten, Allan Jabri, and Nicolas Vasilache. Learning visual features from large weakly supervised data. In ECCV, 2016. 2
  18. 18.Sahar Kazemzadeh, Vicente Ordonez, Mark Matten, and Tamara Berg. Referitgame: Referring to objects in photographs of natural scenes. In EMNLP, 2014. 1
  19. 19.Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization. In ICLR, 2015. 5
  20. 20.Daqing Liu, Hanwang Zhang, Feng Wu, and Zheng-Jun Zha. Learning to assemble neural module tree networks for visual grounding. In ICCV, 2019. 1
  21. 21.Vivek Mittal. Attngrounder: Talking to cars with attention. In ECCV, 2020. 1
  22. 22.Bryan A Plummer, Liwei Wang, Chris M Cervantes, Juan C Caicedo, Julia Hockenmaier, and Svetlana Lazebnik. Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models. In ICCV, 2015. 1
  23. 23.Charles R Qi, Li Yi, Hao Su, and Leonidas J Guibas. Pointnet++: Deep hierarchical feature learning on point sets in a metric space. arXiv, 2017. 4
  24. 24.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. arXiv, 2021. 2
  25. 25.Junha Roh, Karthik Desingh, Ali Farhadi, and Dieter Fox. Languagerefer: Spatial-language model for 3d visual grounding. In CoRL, 2021. 1, 2, 4, 5, 6
  26. 26.Manolis Savva, Abhishek Kadian, Oleksandr Maksymets, Yili Zhao, Erik Wijmans, Bhavana Jain, Julian Straub, Jia Liu, Vladlen Koltun, Jitendra Malik, et al. Habitat: A platform for embodied ai research. In ICCV, 2019. 1
  27. 27.Franco Scarselli, Marco Gori, Ah Chung Tsoi, Markus Hagenbuchner, and Gabriele Monfardini. The graph neural network model. IEEE transactions on neural networks, 2008. 2
  28. 28.Hang Su, Subhransu Maji, Evangelos Kalogerakis, and Erik Learned-Miller. Multi-view convolutional neural networks for 3d shape recognition. In ICCV, 2015. 3
  29. 29.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In NIPS, 2017. 2, 4, 5, 7, 8
  30. 30.Liwei Wang, Jing Huang, Yin Li, Kun Xu, Zhengyuan Yang, and Dong Yu. Improving weakly supervised visual grounding by contrastive knowledge distillation. In CVPR, 2021. 1
  31. 31.Xin Wang, Qiuyuan Huang, Asli Celikyilmaz, Jianfeng Gao, Dinghan Shen, Yuan-Fang Wang, William Yang Wang, and Lei Zhang. Reinforced cross-modal matching and self-supervised imitation learning for vision-language navigation. In CVPR, 2019. 1
  32. 32.Fei Xia, Amir R Zamir, Zhiyang He, Alexander Sax, Jitendra Malik, and Silvio Savarese. Gibson env: Real-world perception for embodied agents. In CVPR, 2018. 1
  33. 33.Zhengyuan Yang, Songyang Zhang, Liwei Wang, and Jiebo Luo. Sat: 2d semantics assisted training for 3d visual grounding. In ICCV, 2021. 1, 2, 4, 5, 6
  34. 34.Licheng Yu, Zhe Lin, Xiaohui Shen, Jimei Yang, Xin Lu, Mohit Bansal, and Tamara L Berg. Mattnet: Modular attention network for referring expression comprehension. In CVPR, 2018. 1
  35. 35.Licheng Yu, Patrick Poirson, Shan Yang, Alexander C Berg, and Tamara L Berg. Modeling context in referring expressions. In ECCV, 2016. 1
  36. 36.Zhihao Yuan, Xu Yan, Yinghong Liao, Ruimao Zhang, Sheng Wang, Zhen Li, and Shuguang Cui. Instancerefer: Cooperative holistic understanding for visual grounding on point clouds through instance multi-level contextual referring. In ICCV, 2021. 1, 2, 5, 6, 7
  37. 37.Yuhao Zhang, Hang Jiang, Yasuhide Miura, Christopher D Manning, and Curtis P Langlotz. Contrastive learning of medical visual representations from paired images and text. arXiv, 2020. 2
  38. 38.Lichen Zhao, Daigang Cai, Lu Sheng, and Dong Xu. 3dvg-transformer: Relation modeling for visual grounding on point clouds. In ICCV, 2021. 2, 6
  39. 39.Fengda Zhu, Yi Zhu, Xiaojun Chang, and Xiaodan Liang. Vision-language navigation with self-supervised auxiliary reasoning tasks. In CVPR, 2020. 1

Citation

MLA
Huang, S., et al. “Multi-View Transformer for 3D Visual Grounding”. arXiv, 2022, http://arxiv.org/abs/2204.02174v1.
APA
Huang, S., Chen, Y., Jia, J., & Wang, L. (2022). Multi-View Transformer for 3D Visual Grounding. arXiv. http://arxiv.org/abs/2204.02174v1
Chicago
Huang, S., Y. Chen, J. Jia, and L. Wang. 2022. “Multi-View Transformer for 3D Visual Grounding”. arXiv. http://arxiv.org/abs/2204.02174v1.
Harvard
Huang, S. et al. (2022) “Multi-View Transformer for 3D Visual Grounding”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2204.02174v1.
Vancouver
1. Huang S, Chen Y, Jia J, Wang L (2022) Multi-View Transformer for 3D Visual Grounding. arXiv

BibTeX

@article{huang2022multi,
  title = {Multi-View Transformer for 3D Visual Grounding},
  author = {Huang, Shijia and Chen, Yilun and Jia, Jiaya and Wang, Liwei},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2204.02174v1},
  eprint = {2204.02174}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE