Improving Visual Grounding with Visual-Linguistic Verification and Iterative Reasoning

Li YangYan XuChunfeng YuanWei LiuBing LiWeiming Hu

article2022CVPR187 citations

Presents a transformer-based visual grounding framework that integrates visual-linguistic verification, language-guided context encoding, and iterative multi-stage decoding to localize referred objects directly without relying on predefined proposals or anchor boxes.

Listen

Visual grounding involves locating specific objects within images using natural language descriptions, serving as a critical bridge between computer vision and language processing. Conventional methods typically treat this as a ranking problem across predefined candidate regions or anchor boxes. However, this traditional approach often misses fine-grained visual contexts and linguistic cues, creating performance bottlenecks whenever initial object proposals are inaccurate or incomplete.

The article demonstrates a dedicated transformer-based framework designed to directly retrieve and localize target objects by generating discriminative, text-guided visual representations and performing multi-stage cross-modal reasoning. To achieve this, the authors designed a visual-linguistic verification module to focus visual features on text-relevant regions while suppressing distractions, a language-guided context encoder to aggregate surrounding spatial and relational context, and an iterative cross-modal decoder that refines target object queries across multiple stages. The framework was evaluated across five standard benchmark datasets: RefCOCO, RefCOCO+, RefCOCOg, ReferItGame, and Flickr30k Entities.

The findings show that the proposed method establishes a new state of the art across all evaluated benchmarks. Compared to leading proposal-based methods, it delivers absolute accuracy gains of up to 4.45% on RefCOCO, 5.94% on RefCOCO+, and 5.49% on RefCOCOg. Against modern one-stage approaches, it improves accuracy on RefCOCO validation and test sets by 5.10 and 6.34 percentage points, while also outperforming prior transformer-based methods like TransVG by 2.14% to 9.37% across the RefCOCO benchmark family. Ablation experiments confirm that incorporating iterative decoder stages, context encoding, and verification progressively raises accuracy from a 63.64% baseline to 71.62%, adding only an 8.81 million parameter increase (about 6.14%) and a 1.68% rise in computational complexity.

These results indicate that directly learning text-conditioned visual features and iteratively reasoning over multimodal cues significantly outperforms both rigid proposal-ranking pipelines and generic transformer fusion architectures. By eliminating reliance on separate object proposal generators, the architecture streamlines training and improves precision without imposing excessive computational overhead. However, the article notes a limitation: the system was trained exclusively on standard visual grounding datasets with restricted vocabulary sizes, which may constrain generalization to broader, open-domain language queries. The authors recommend expanding training to larger-scale datasets to enhance real-world generalization before deploying in open-domain applications.

arXiv: 2205.00272
Cover for Improving Visual Grounding with Visual-Linguistic Verification and Iterative Reasoning

Abstract

Visual grounding is a task to locate the target indicated by a natural language expression. Existing methods extend the generic object detection framework to this problem. They base the visual grounding on the features from pre-generated proposals or anchors, and fuse these features with the text embeddings to locate the target mentioned by the text. However, modeling the visual features from these predefined locations may fail to fully exploit the visual content and attribute information in the text query, which limits their performance. In this paper, we propose a transformer-based framework for accurate visual grounding by establishing text-conditioned discriminative features and performing multi-stage cross-modal reasoning. Specifically, we develop a visual-linguistic verification module to focus the visual features on regions relevant to the textual descriptions while suppressing the unrelated areas. A language-guided feature encoder is also devised to aggregate the visual contexts of the target object to improve the object's distinctiveness. To retrieve the target from the encoded visual features, we further propose a multi-stage cross-modal decoder to iteratively speculate on the correlations between the image and text for accurate target localization. Extensive experiments on five widely used datasets validate the efficacy of our proposed components and demonstrate state-of-the-art performance.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Visual Grounding
  • 2.2 Visual Transformer
  • 3 Method
  • 3.1 The Overall Network
  • 3.2 Visual-Linguistic Verification Module
  • 3.3 Language-guided Context Encoder
  • 3.4 Multi-stage Cross-modal Decoder
  • 3.5 Training Loss
  • 4 Experiments
  • 4.1 Datasets
  • 4.2 Implementation Details
  • 4.3 Comparisons with State-of-the-art Methods
  • 4.4 Ablation Study
  • 4.4.1 The Component Modules
  • 4.4.2 The Decoder Stages
  • 4.4.3 Visual-Linguistic Feature Learning
  • 4.5 Visualization
  • 5 Conclusions and Limitations
  • References

Knowls

  1. Knowl 1 — VLTVG Framework for Visual Grounding

    model/method

    The Visual-Linguistic Verification and Iterative Reasoning (VLTVG) framework directly retrieves target object representations for visual grounding by directly predicting bounding box coordinates rather than ranking pre-generated proposals or anchor points.

    The framework consists of four core components:

    1. Feature Extraction: Given an image and a natural language expression, visual features Fv∈RC×H×WF_v \in \mathbb{R}^{C \times H \times W} are generated using a CNN backbone (ResNet-50 or ResNet-101) followed by a 6-layer Transformer encoder (initialized from DETR), while text embeddings Fl∈RC×LF_l \in \mathbb{R}^{C \times L} are generated via BERT, where CC is the channel dimension, H×WH \times W is the spatial resolution, and LL is the token sequence length.
    2. Visual-Linguistic Verification Module: Computes fine-grained semantic relevance scores S∈RH×WS \in \mathbb{R}^{H \times W} across spatial coordinates to highlight text-relevant visual regions and suppress irrelevant backgrounds.
    3. Language-Guided Context Encoder: Leverages text semantics to guide self-attention over visual tokens, extracting relational and positional context features Fvc∈RC×H×WF_{vc} \in \mathbb{R}^{C \times H \times W} to form modulated discriminative features F^v=(Fv+Fvc)⋅S\hat{F}_v = (F_v + F_{vc}) \cdot S.
    4. Multi-Stage Cross-Modal Decoder: A stack of NN cross-modal reasoning stages that iteratively update a learnable target query tqt_q by attending alternately to textual embeddings and discriminative visual features to regress the target bounding box.
  2. Knowl 2 — Visual-Linguistic Verification Module

    model/method

    The Visual-Linguistic Verification module computes spatial correlation scores between visual features and text expressions to suppress regions irrelevant to the query.

    First, a multi-head cross-attention layer treats the visual feature map Fv∈RC×H×WF_v \in \mathbb{R}^{C \times H \times W} as queries and the textual embeddings Fl∈RC×LF_l \in \mathbb{R}^{C \times L} as keys and values, producing a spatially aligned linguistic semantic map Fs∈RC×H×WF_s \in \mathbb{R}^{C \times H \times W}.

    Next, FvF_v and FsF_s are projected into a common semantic space via linear projection and L2L_2-normalization, yielding Fv′F'_v and Fs′F'_s. The verification score S(x,y)∈RS(x, y) \in \mathbb{R} at spatial location (x,y)(x, y) is computed as:

    S(x,y)=α⋅exp⁡(−(1−Fv′(x,y)TFs′(x,y))22σ2)S(x, y) = \alpha \cdot \exp\left(-\frac{\left(1 - {F'_v(x, y)}^T F'_s(x, y)\right)^2}{2\sigma^2}\right)

    where α\alpha and σ\sigma are learnable parameters (initialized to α=1.0\alpha = 1.0 and σ=0.5\sigma = 0.5). Modulating the visual features pixel-wise via F^v=Fv⋅S\hat{F}_v = F_v \cdot S acts as a soft spatial mask that dampens unrelated image regions before cross-modal reasoning.

  3. Knowl 3 — Language-Guided Context Encoder

    model/method

    The Language-Guided Context Encoder captures visual interaction and relative positional contexts of referred objects using linguistic guidance.

    1. Text Semantic Gathering: Visual feature map Fv∈RC×H×WF_v \in \mathbb{R}^{C \times H \times W} queries the textual embeddings Fl∈RC×LF_l \in \mathbb{R}^{C \times L} through multi-head cross-attention to produce text representation Fc∈RC×H×WF_c \in \mathbb{R}^{C \times H \times W} aligned with visual locations.
    2. Language-Guided Self-Attention: The combined representation Fv+FcF_v + F_c acts as both query and key in a multi-head self-attention layer with relative positional encodings. For attention head kk with linear projections WQW_Q and WKW_K, channel dimension dkd_k, and sinusoidal relative positional encodings R(i−j)R(i - j), the attention weight between positions ii and jj is:

    attni,j=softmax(Q(i)T(K(j)+WKTR(i−j))dk)\text{attn}_{i,j} = \text{softmax}\left(\frac{Q(i)^T\left(K(j) + W_K^T R(i - j)\right)}{\sqrt{d_k}}\right)

    where Q=WQT(Fv+Fc)Q = W_Q^T (F_v + F_c) and K=WKT(Fv+Fc)K = W_K^T (F_v + F_c).

    The resulting visual context map Fvc∈RC×H×WF_{vc} \in \mathbb{R}^{C \times H \times W} is fused with FvF_v and the visual-linguistic verification score map SS to form discriminative feature representation:

    F^v=(Fv+Fvc)⋅S\hat{F}_v = (F_v + F_{vc}) \cdot S

  4. Knowl 4 — Multi-Stage Cross-Modal Decoder

    model/method

    The Multi-Stage Cross-Modal Decoder performs iterative cross-modal reasoning over NN stages (default N=6N=6, with unshared weights across stages) to refine a learnable target query tq1∈RC×1t_q^1 \in \mathbb{R}^{C \times 1} into a precise referred object representation.

    At each stage i∈{1,…,N}i \in \{1, \dots, N\}:

    1. Text Querying: The query tqit_q^i attends to textual embeddings Fl∈RC×LF_l \in \mathbb{R}^{C \times L} using multi-head attention to extract semantic descriptions tl∈RC×1t_l \in \mathbb{R}^{C \times 1}.
    2. Visual Feature Gathering: Using tlt_l as query and the discriminative visual feature F^v∈RC×H×W\hat{F}_v \in \mathbb{R}^{C \times H \times W} as key/query guidance, visual features of interest are gathered from the original visual feature map Fv∈RC×H×WF_v \in \mathbb{R}^{C \times H \times W} (serving as value) via multi-head attention to produce visual representation tv∈RC×1t_v \in \mathbb{R}^{C \times 1}.
    3. Query Update:

    tq′=LN(tqi+tv)t'_q = \text{LN}(t_q^i + t_v) tqi+1=LN(tq′+FFN(tq′))t_q^{i+1} = \text{LN}(t'_q + \text{FFN}(t'_q))

    where LN(⋅)\text{LN}(\cdot) is layer normalization and FFN(⋅)\text{FFN}(\cdot) is a two-layer feed-forward network with ReLU activation.

    At each stage ii, a 3-layer MLP with ReLU activation regresses the bounding box b^i∈R4\hat{b}^i \in \mathbb{R}^4 directly from tqi+1t_q^{i+1}.

  5. Knowl 5 — Multi-Stage Direct Bounding Box Regression Loss

    equation

    The VLTVG model is trained end-to-end by supervising bounding box predictions across all NN stages of the cross-modal decoder directly against the ground-truth bounding box bb without positive/negative sample assignment:

    L=∑i=1N(λgiouLgiou(b,b^i)+λL1LL1(b,b^i))\mathcal{L} = \sum_{i=1}^N \left( \lambda_{\text{giou}} \mathcal{L}_{\text{giou}}(b, \hat{b}^i) + \lambda_{L1} \mathcal{L}_{L1}(b, \hat{b}^i) \right)

    where:

    • NN is the number of decoder reasoning stages (N=6N = 6).
    • b^i\hat{b}^i is the predicted 4-dimensional bounding box coordinates at stage ii.
    • bb is the ground-truth bounding box.
    • Lgiou(⋅,⋅)\mathcal{L}_{\text{giou}}(\cdot, \cdot) is the Generalized Intersection over Union (GIoU) loss.
    • LL1(⋅,⋅)\mathcal{L}_{L1}(\cdot, \cdot) is the standard L1L_1 regression loss.
    • λgiou=2\lambda_{\text{giou}} = 2 and λL1=5\lambda_{L1} = 5 are balancing hyperparameters.
  6. Knowl 6 — Performance on RefCOCO, RefCOCO+, and RefCOCOg Benchmarks

    data/table

    VLTVG outperforms two-stage, one-stage, and previous transformer-based methods on RefCOCO, RefCOCO+, and RefCOCOg datasets (evaluated by accuracy at IoU>0.5\text{IoU} > 0.5):

    Models Venue Backbone RefCOCO RefCOCO+ RefCOCOg
    val testA testB val testA testB val-g val-u test-u
    Two-stage:
    CMN CVPR'17 VGG16 - 71.03 65.77 - 54.32 47.76 57.47 - -
    VC CVPR'18 VGG16 - 73.33 67.44 - 58.40 53.18 62.30 - -
    ParalAttn CVPR'18 VGG16 - 75.31 65.52 - 61.34 50.86 58.03 - -
    MAttNet CVPR'18 ResNet-101 76.65 81.14 69.99 65.33 71.62 56.02 - 66.58 67.27
    LGRANs CVPR'19 VGG16 - 76.60 66.40 - 64.00 53.40 61.78 - -
    DGA ICCV'19 VGG16 - 78.42 65.53 - 69.07 51.99 - - 63.28
    RvG-Tree TPAMI'19 ResNet-101 75.06 78.61 69.85 63.51 67.45 56.66 - 66.95 66.51
    NMTree ICCV'19 ResNet-101 76.41 81.21 70.09 66.46 72.02 57.52 64.62 65.87 66.44
    Ref-NMS AAAI'21 ResNet-101 80.70 84.00 76.04 68.25 73.68 59.42 - 70.55 70.62
    One-stage:
    SSG arXiv'18 DarkNet-53 - 76.51 67.50 - 62.14 49.27 47.47 58.80 -
    FAOA ICCV'19 DarkNet-53 72.54 74.35 68.50 56.81 60.23 49.60 56.12 61.33 60.36
    RCCF CVPR'20 DLA-34 - 81.06 71.85 - 70.35 56.32 - - 65.73
    ReSC-Large ECCV'20 DarkNet-53 77.63 80.45 72.30 63.59 68.36 56.81 63.12 67.30 67.20
    LBYL-Net CVPR'21 DarkNet-53 79.67 82.91 74.15 68.64 73.38 59.49 62.70 - -
    Transformer:
    TransVG ICCV'21 ResNet-50 80.32 82.67 78.12 63.50 68.15 55.63 66.56 67.66 67.44
    TransVG ICCV'21 ResNet-101 81.02 82.72 78.35 64.82 70.70 56.94 67.02 68.67 67.73
    VLTVG (ours) - ResNet-50 84.53 87.69 79.22 73.60 78.37 64.53 72.53 74.90 73.88
    VLTVG (ours) - ResNet-101 84.77 87.24 80.49 74.19 78.93 65.17 72.98 76.04 74.18

    Compared to TransVG with ResNet-101, VLTVG achieves absolute accuracy gains of 2.14% to 4.52% on RefCOCO, 8.23% to 9.37% on RefCOCO+, and 5.96% to 7.37% on RefCOCOg.

  7. Knowl 7 — Performance on ReferItGame and Flickr30k Entities Benchmarks

    data/table

    VLTVG establishes new state-of-the-art results on the test sets of ReferItGame and Flickr30k Entities (accuracy at IoU>0.5\text{IoU} > 0.5):

    Models Backbone ReferItGame test Flickr30k test
    Two-stage:
    CMN VGG16 28.33 -
    VC VGG16 31.13 -
    MAttNet ResNet-101 29.04 -
    Similarity Net ResNet-101 34.54 60.89
    CITE ResNet-101 35.07 61.33
    DDPN ResNet-101 63.00 73.30
    One-stage:
    SSG DarkNet-53 54.24 -
    ZSGNet ResNet-50 58.63 63.39
    FAOA DarkNet-53 60.67 68.71
    RCCF DLA-34 63.79 -
    ReSC-Large DarkNet-53 64.60 69.28
    LBYL-Net DarkNet-53 67.47 -
    Transformer-based:
    TransVG ResNet-50 69.76 78.47
    TransVG ResNet-101 70.73 79.10
    VLTVG (ours) ResNet-50 71.60 79.18
    VLTVG (ours) ResNet-101 71.98 79.84

    VLTVG achieves 71.98% on ReferItGame test (+1.25% over TransVG) and 79.84% on Flickr30k Entities test (+0.74% over TransVG). The smaller gain on Flickr30k Entities is attributed to its expressions consisting mostly of short noun phrases with fewer relational visual contexts.

  8. Knowl 8 — Ablation Analysis of VLTVG Architectural Components and Feature Fusion

    data/table

    Ablation experiments on RefCOCOg (val-g) validate each module's contribution and compare the proposed fusion against stacking standard transformer encoder layers:

    Configuration Multi-stage Decoder Context Encoder V-L Verification #params Acc (%)
    Baseline (single-stage decoder) 143.37M 63.64
    + Multi-stage Decoder (N=6N=6) ✓ 151.26M 66.02
    + Language-guided Context Encoder ✓ ✓ 151.79M 68.44
    Full VLTVG ✓ ✓ ✓ 152.18M 71.62
    V-L Feature Learning Strategy #params GFLOPS Acc (%)
    None 151.26M 41.39 66.02
    Trans. encoder layers (×1\times 1) 152.69M 42.08 69.37
    Trans. encoder layers (×2\times 2) 154.00M 42.77 69.22
    Trans. encoder layers (×3\times 3) 155.32M 43.46 69.15
    Trans. encoder layers (×4\times 4) 156.63M 44.15 69.55
    Ours (V-L verification + context) 152.18M 41.79 71.62

    Key observations:

    1. The Multi-Stage Decoder (+2.38%), Context Encoder (+2.42%), and Verification module (+3.18%) contribute cumulatively to a +7.98% total improvement over the baseline while adding only 8.81M parameters (+6.14%) and 0.69 GFLOPS (+1.68%).
    2. Stacking 1 to 4 standard transformer encoder layers plateaus around 69.55% accuracy, whereas VLTVG reaches 71.62% (+2.07% over best stacked encoder) using fewer parameters and lower compute.
  9. Knowl 9 — Ablation on Decoder Stages Count

    data/table

    Evaluating the multi-stage cross-modal decoder on RefCOCOg (val-g) shows the effect of varying stage count NN:

    Decoder Stages (NN) #params GFLOPS Acc (%)
    N=1N = 1 143.37M 41.10 63.64
    N=2N = 2 144.95M 41.15 65.05
    N=4N = 4 148.10M 41.27 65.70
    N=6N = 6 151.26M 41.39 66.02
    N=8N = 8 154.42M 41.51 65.97

    Accuracy increases steadily up to N=6N = 6 (+2.38% over N=1N=1), after which performance plateaus (65.97% at N=8N=8). Increasing from N=1N=1 to N=6N=6 adds only 5.5% parameters and 0.71% GFLOPS, confirming N=6N=6 as the optimal trade-off.

  10. Knowl 10 — Corpus Size and Lexical Generalization Limitations

    limitation

    A primary limitation of VLTVG is that it is trained exclusively on visual grounding datasets containing closed corpora with limited linguistic variety. Consequently, the model may experience degraded localization performance and limited generalization when faced with open-vocabulary, highly descriptive, or unconstrained out-of-distribution natural language expressions.

Coverage note — None was omitted; all key architectural modules, equations, loss definitions, experimental benchmark comparisons (Table 1, Table 2), ablation studies (Table 3, Table 4, Table 5), and stated limitations are fully represented.

References

  1. 1.Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. End-to-end object detection with transformers. In European Conference on Computer Vision, pages 213–229. Springer, 2020. 2, 5
  2. 2.Long Chen, Wenbo Ma, Jun Xiao, Hanwang Zhang, and Shih-Fu Chang. Ref-nms: Breaking proposal bottlenecks in two-stage referring expression grounding. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, pages 1036–1044, 2021. 6
  3. 3.Xinpeng Chen, Lin Ma, Jingyuan Chen, Zequn Jie, Wei Liu, and Jiebo Luo. Real-time referring expression comprehension by single-stage grounding network. arXiv preprint arXiv:1812.03426, 2018. 1, 2, 6, 7
  4. 4.Yanbei Chen, Shaogang Gong, and Loris Bazzani. Image search with text feedback by visiolinguistic attention learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3001–3011, 2020. 3
  5. 5.Jiajun Deng, Zhengyuan Yang, Tianlang Chen, Wengang Zhou, and Houqiang Li. Transvg: End-to-end visual grounding with transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 1769–1779, 2021. 2, 5, 6, 7
  6. 6.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018. 2, 3, 5
  7. 7.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. In International Conference on Learning Representations, 2020. 2, 3
  8. 8.Hugo Jair Escalante, Carlos A Hernandez, Jesus A Gonzalez, ´ Aurelio Lopez-L ´ opez, Manuel Montes, Eduardo F Morales, ´ L Enrique Sucar, Luis Villasenor, and Michael Grubinger. The segmented and annotated iapr tc-12 benchmark. Computer vision and image understanding, 114(4):419–428, 2010. 5
  9. 9.Xavier Glorot and Yoshua Bengio. Understanding the difficulty of training deep feedforward neural networks. In Proceedings of the thirteenth international conference on artificial intelligence and statistics, pages 249–256, 2010. 5
  10. 10.Kaiming He, Georgia Gkioxari, Piotr Dollar, and Ross Gir- ´ shick. Mask r-cnn. In Proceedings of the IEEE international conference on computer vision, pages 2961–2969, 2017. 2, 6
  11. 11.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016. 2, 3
  12. 12.Richang Hong, Daqing Liu, Xiaoyu Mo, Xiangnan He, and Hanwang Zhang. Learning to compose and reason with language tree structures for visual grounding. IEEE transactions on pattern analysis and machine intelligence, 2019. 2, 6
  13. 13.Ronghang Hu, Marcus Rohrbach, Jacob Andreas, Trevor Darrell, and Kate Saenko. Modeling relationships in referential expressions with compositional modular networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 1115–1124, 2017. 2, 6, 7
  14. 14.Ronghang Hu, Huazhe Xu, Marcus Rohrbach, Jiashi Feng, Kate Saenko, and Trevor Darrell. Natural language object retrieval. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 4555–4564, 2016. 5
  15. 15.Binbin Huang, Dongze Lian, Weixin Luo, and Shenghua Gao. Look before you leap: Learning landmark features for one-stage visual grounding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 16888–16897, 2021. 6, 7
  16. 16.Sahar Kazemzadeh, Vicente Ordonez, Mark Matten, and Tamara Berg. Referitgame: Referring to objects in photographs of natural scenes. In Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP), pages 787–798, 2014. 2, 5, 6, 7
  17. 17.Kuang-Huei Lee, Xi Chen, Gang Hua, Houdong Hu, and Xiaodong He. Stacked cross attention for image-text matching. In Proceedings of the European Conference on Computer Vision (ECCV), pages 201–216, 2018. 3
  18. 18.Yue Liao, Si Liu, Guanbin Li, Fei Wang, Yanjie Chen, Chen Qian, and Bo Li. A real-time cross-modality correlation filtering method for referring expression comprehension. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10880–10889, 2020. 1, 2, 5, 6, 7
  19. 19.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollar, and C Lawrence ´ Zitnick. Microsoft coco: Common objects in context. In European Conference on Computer Vision, pages 740–755. Springer, 2014. 5
  20. 20.Daqing Liu, Hanwang Zhang, Feng Wu, and Zheng-Jun Zha. Learning to assemble neural module tree networks for visual grounding. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 4673–4682, 2019. 2, 6
  21. 21.Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. In International Conference on Learning Representations, 2018. 5
  22. 22.Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks. Advances in neural information processing systems, 32, 2019. 3
  23. 23.Junhua Mao, Jonathan Huang, Alexander Toshev, Oana Camburu, Alan L Yuille, and Kevin Murphy. Generation and comprehension of unambiguous object descriptions. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 11–20, 2016. 1, 2, 5, 6
  24. 24.Varun K Nagaraja, Vlad I Morariu, and Larry S Davis. Modeling context between objects for referring expression understanding. In European Conference on Computer Vision, pages 792–807. Springer, 2016. 1, 5
  25. 25.Bryan A Plummer, Paige Kordas, M Hadi Kiapour, Shuai Zheng, Robinson Piramuthu, and Svetlana Lazebnik. Conditional image-text embedding networks. In Proceedings of the European Conference on Computer Vision (ECCV), pages 249–264, 2018. 7
  26. 26.Bryan A Plummer, Liwei Wang, Chris M Cervantes, Juan C Caicedo, Julia Hockenmaier, and Svetlana Lazebnik. Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models. In Proceedings of the IEEE international conference on computer vision, pages 2641–2649, 2015. 2, 5, 6, 7
  27. 27.Joseph Redmon and Ali Farhadi. Yolov3: An incremental improvement. arXiv preprint arXiv:1804.02767, 2018. 2
  28. 28.Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. Faster r-cnn: Towards real-time object detection with region proposal networks. IEEE Transactions on Pattern Analysis and Machine Intelligence, 39(6):1137–1149, 2016. 2
  29. 29.Hamid Rezatofighi, Nathan Tsoi, JunYoung Gwak, Amir Sadeghian, Ian Reid, and Silvio Savarese. Generalized intersection over union: A metric and a loss for bounding box regression. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 658–666, 2019. 5
  30. 30.Arka Sadhu, Kan Chen, and Ram Nevatia. Zero-shot grounding of objects from natural language queries. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 4694–4703, 2019. 7
  31. 31.Rui Su, Qian Yu, and Dong Xu. Stvgbert: A visual-linguistic transformer based framework for spatio-temporal video grounding. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 1533–1542, 2021. 3
  32. 32.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Advances in neural information processing systems, pages 5998–6008, 2017. 2, 4
  33. 33.Liwei Wang, Yin Li, Jing Huang, and Svetlana Lazebnik. Learning two-branch neural networks for image-text matching tasks. IEEE Transactions on Pattern Analysis and Machine Intelligence, 41(2):394–407, 2018. 1, 2, 7
  34. 34.Peng Wang, Qi Wu, Jiewei Cao, Chunhua Shen, Lianli Gao, and Anton van den Hengel. Neighbourhood watch: Referring expression comprehension via language-guided graph attention networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1960–1968, 2019. 2, 6
  35. 35.Yan Xu, Zhaoyang Huang, Kwan-Yee Lin, Xinge Zhu, Jianping Shi, Hujun Bao, Guofeng Zhang, and Hongsheng Li. Selfvoxelo: Self-supervised lidar odometry with voxel-based deep neural networks. In Conference on Robot Learning, pages 115–125. PMLR, 2021. 2
  36. 36.Yan Xu, Junyi Lin, Jianping Shi, Guofeng Zhang, Xiaogang Wang, and Hongsheng Li. Robust self-supervised lidar odometry via representative structure discovery and 3d inherent error modeling. IEEE Robotics and Automation Letters, 2022. 2
  37. 37.Yan Xu, Kwan-Yee Lin, Guofeng Zhang, Xiaogang Wang, and Hongsheng Li. Rnnpose: Recurrent 6-dof object pose refinement with robust correspondence field estimation and pose optimization, 2022. 2
  38. 38.Yan Xu, Xinge Zhu, Jianping Shi, Guofeng Zhang, Hujun Bao, and Hongsheng Li. Depth completion from sparse lidar data with depth-normal constraints. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 2811–2820, 2019. 2
  39. 39.Li Yang, Yan Xu, Shaoru Wang, Chunfeng Yuan, Ziqi Zhang, Bing Li, and Weiming Hu. Pdnet: Towards better one-stage object detection with prediction decoupling. arXiv preprint arXiv:2104.13876, 2021. 2
  40. 40.Sibei Yang, Guanbin Li, and Yizhou Yu. Dynamic graph attention for referring expression comprehension. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 4644–4653, 2019. 2, 6
  41. 41.Sibei Yang, Guanbin Li, and Yizhou Yu. Graph-structured referring expression reasoning in the wild. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9952–9961, 2020. 2
  42. 42.Zhengyuan Yang, Tianlang Chen, Liwei Wang, and Jiebo Luo. Improving one-stage visual grounding by recursive subquery construction. In European Conference on Computer Vision, pages 387–404. Springer, 2020. 2, 5, 6, 7
  43. 43.Zhengyuan Yang, Boqing Gong, Liwei Wang, Wenbing Huang, Dong Yu, and Jiebo Luo. A fast and accurate one-stage approach to visual grounding. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 4683–4693, 2019. 1, 2, 5, 6, 7
  44. 44.Licheng Yu, Zhe Lin, Xiaohui Shen, Jimei Yang, Xin Lu, Mohit Bansal, and Tamara L Berg. Mattnet: Modular attention network for referring expression comprehension. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 1307–1315, 2018. 2, 5, 6, 7
  45. 45.Licheng Yu, Patrick Poirson, Shan Yang, Alexander C Berg, and Tamara L Berg. Modeling context in referring expressions. In European Conference on Computer Vision, pages 69–85. Springer, 2016. 1, 2, 5, 6
  46. 46.Zhou Yu, Jun Yu, Chenchao Xiang, Zhou Zhao, Qi Tian, and Dacheng Tao. Rethinking diversified and discriminative proposal generation for visual grounding. International Joint Conference on Artificial Intelligence (IJCAI), 2018. 7
  47. 47.Hanwang Zhang, Yulei Niu, and Shih-Fu Chang. Grounding referring expressions in images by variational context. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 4158–4166, 2018. 2, 6, 7
  48. 48.Xinge Zhu, Yuexin Ma, Tai Wang, Yan Xu, Jianping Shi, and Dahua Lin. Ssn: Shape signature networks for multi-class object detection from point clouds. In European Conference on Computer Vision, pages 581–597. Springer, 2020. 2
  49. 49.Xizhou Zhu, Weijie Su, Lewei Lu, Bin Li, Xiaogang Wang, and Jifeng Dai. Deformable detr: Deformable transformers for end-to-end object detection. In International Conference on Learning Representations, 2020. 2, 3
  50. 50.Bohan Zhuang, Qi Wu, Chunhua Shen, Ian Reid, and Anton Van Den Hengel. Parallel attention: A unified framework for visual object discovery through dialogs and queries. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 4252–4261, 2018. 2, 6

Citation

MLA
Yang, L., et al. “Improving Visual Grounding with Visual-Linguistic Verification and Iterative Reasoning”. arXiv, 2022, http://arxiv.org/abs/2205.00272v2.
APA
Yang, L., Xu, Y., Yuan, C., Liu, W., Li, B., & Hu, W. (2022). Improving Visual Grounding with Visual-Linguistic Verification and Iterative Reasoning. arXiv. http://arxiv.org/abs/2205.00272v2
Chicago
Yang, L., Y. Xu, C. Yuan, W. Liu, B. Li, and W. Hu. 2022. “Improving Visual Grounding with Visual-Linguistic Verification and Iterative Reasoning”. arXiv. http://arxiv.org/abs/2205.00272v2.
Harvard
Yang, L. et al. (2022) “Improving Visual Grounding with Visual-Linguistic Verification and Iterative Reasoning”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2205.00272v2.
Vancouver
1. Yang L, Xu Y, Yuan C, Liu W, Li B, Hu W (2022) Improving Visual Grounding with Visual-Linguistic Verification and Iterative Reasoning. arXiv

BibTeX

@article{yang2022improving,
  title = {Improving Visual Grounding with Visual-Linguistic Verification and Iterative Reasoning},
  author = {Yang, Li and Xu, Yan and Yuan, Chunfeng and Liu, Wei and Li, Bing and Hu, Weiming},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2205.00272v2},
  eprint = {2205.00272}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE