LAVT: Language-Aware Vision Transformer for Referring Image Segmentation

Zhao YangJiaqi WangYansong TangKai ChenHengshuang ZhaoPhilip H. S. Torr

article2022CVPR561 citations

Proposes an early fusion framework that integrates linguistic cues directly into the intermediate layers of a vision Transformer encoder, replacing complex decoders with a lightweight predictor to set new performance standards on referring image segmentation benchmarks.

Listen

Referring image segmentation requires an artificial intelligence system to identify and outline a specific object in an image based on a natural language description. This capability is critical for practical applications such as language-guided robotics and automated image editing. Conventional approaches extract image and text features independently and combine them afterward using complex, computationally heavy multi-modal decoders. However, these late-fusion architectures miss valuable opportunities to let language guide visual processing from the start, leaving performance sub-optimal.

The article demonstrates that directly integrating linguistic features into the visual encoding stages of a vision Transformer architecture yields superior multi-modal alignment. It evaluates this new architecture, termed the Language-Aware Vision Transformer (LAVT), to show that early cross-modal feature fusion makes complicated decoders obsolete while setting a new performance standard for the task.

The researchers designed an architecture that pairs a BERT language model with a multi-stage hierarchical vision backbone (Swin Transformer). Rather than waiting until visual extraction finishes, the model fuses text embeddings into visual features at every encoding stage using a lightweight pixel-word attention module and a gating mechanism that regulates language flow. The resulting language-enriched visual features are then processed by a simple mask prediction head. The approach was systematically trained and tested on three standard benchmark datasets: RefCOCO, RefCOCO+, and G-Ref (comprising tens of thousands of images and complex text expressions).

The evaluation produced four primary findings. First, the proposed architecture established new state-of-the-art results across all evaluated benchmarks, achieving overall Intersection-over-Union (IoU) scores of 72.73% on RefCOCO, 62.14% on RefCOCO+, and 61.24% on G-Ref (UMD partition). Second, these results surpassed existing leading methods by substantial margins, improving overall IoU by 5.4 to 9.2 percentage points across various test splits without requiring extra pre-training data. Third, ablation experiments proved that the multi-stage language pathway and dense pixel-word attention are vital, as removing them caused overall IoU to drop by approximately 1.7 to 2.0 percentage points. Finally, experiments revealed that adding a heavy Transformer decoder on top of this early-fusion encoder yielded virtually no additional accuracy gain (only a 0.11% increase in low-threshold precision), confirming that early integration captures the necessary alignment.

These findings indicate that early cross-modal alignment within the visual backbone is fundamentally more effective than post-extraction fusion. For development teams, this allows systems to replace cumbersome cross-modal decoders with simpler, lighter mask predictors, reducing pipeline complexity while improving accuracy on both simple and complex descriptive queries. This architectural shift challenges the prevailing late-fusion paradigm and offers a more scalable framework for vision-language systems.

Organizations developing vision-language applications should adopt early-fusion Transformer designs for dense prediction tasks and retire multi-stage pipelines that rely on heavy cross-modal decoders. Teams should also audit their evaluation datasets, as real-world deployment requires addressing low-quality or ambiguous annotations. Future work should focus on extending this unified early-fusion framework to video segmentation and other multi-modal interaction tasks.

Confidence in these findings is high due to rigorous benchmarking across multiple established datasets and consistent gains across diverse metric thresholds. Nonetheless, stakeholders should note two limitations: the method relies on large, pre-trained backbone models (Swin-B and BERT), and benchmark datasets such as RefCOCO contain instances of ambiguous or noisy language annotations that may affect real-world robustness.

arXiv: 2112.02244
Cover for LAVT: Language-Aware Vision Transformer for Referring Image Segmentation

Abstract

Referring image segmentation is a fundamental vision-language task that aims to segment out an object referred to by a natural language expression from an image. One of the key challenges behind this task is leveraging the referring expression for highlighting relevant positions in the image. A paradigm for tackling this problem is to leverage a powerful vision-language (“cross-modal”) decoder to fuse features independently extracted from a vision encoder and a language encoder. Recent methods have made remarkable advancements in this paradigm by exploiting Transformers as cross-modal decoders, concurrent to the Transformer’s overwhelming success in many other vision-language tasks. Adopting a different approach in this work, we show that significantly better cross-modal alignments can be achieved through the early fusion of linguistic and visual features in intermediate layers of a vision Transformer encoder network. By conducting cross-modal feature fusion in the visual feature encoding stage, we can leverage the well-proven correlation modeling power of a Transformer encoder for excavating helpful multi-modal context. This way, accurate segmentation results are readily harvested with a light-weight mask predictor. Without bells and whistles, our method surpasses the previous state-of-the-art methods on RefCOCO, RefCOCO+, and G-Ref by large margins.

Table of Contents

  • 1. Introduction
  • 2. Related work
  • 3. Method
  • 3.1. Language-aware visual encoding
  • 3.2. Pixel-word attention module
  • 3.3. Language pathway
  • 3.4. Segmentation
  • 3.5. Implementation
  • 4. Experiments
  • 4.1. Datasets and metrics
  • 4.2. Comparison with others
  • 4.3. Ablation study
  • 5. Conclusion
  • References

Knowls

  1. Knowl 1 — Language-Aware Vision Transformer (LAVT) Framework

    model/method

    The Language-Aware Vision Transformer (LAVT) is a framework for referring image segmentation that performs early cross-modal feature fusion within the intermediate stages of a hierarchical vision Transformer encoder (such as Swin Transformer), rather than relying on an isolated visual encoder followed by a heavy cross-modal Transformer decoder.

    Given an input image and a referring text expression of TT words, language features L∈RCt×TL \in \mathbb{R}^{C_t \times T} (Ct=768C_t = 768) are extracted via a base BERT encoder. The vision backbone is organized into four hierarchical stages i∈{1,2,3,4}i \in \{1, 2, 3, 4\}, where each stage comprises a stack of Transformer layers ϕi\phi_i, a Pixel-Word Attention Module (PWAM) θi\theta_i, and a Language Gate (LG) ψi\psi_i:

    1. Stage Transformer layers ϕi\phi_i process features from the previous stage to output visual representations Vi∈RCi×Hi×WiV_i \in \mathbb{R}^{C_i \times H_i \times W_i}.
    2. PWAM θi\theta_i fuses visual features ViV_i with linguistic features LL to produce cross-modal features Fi∈RCi×Hi×WiF_i \in \mathbb{R}^{C_i \times H_i \times W_i}.
    3. The Language Pathway scales FiF_i via the learnable gating unit ψi\psi_i and adds it residually to ViV_i, generating enhanced language-aware visual features Ei∈RCi×Hi×WiE_i \in \mathbb{R}^{C_i \times H_i \times W_i} that serve as input to the next stage ϕi+1\phi_{i+1} (for i∈{1,2,3}i \in \{1, 2, 3\}).
    4. A lightweight top-down segmentation decoder aggregates the multi-modal stage outputs {F1,F2,F3,F4}\{F_1, F_2, F_3, F_4\} to produce pixel-level binary segmentation predictions.
  2. Knowl 2 — Pixel-Word Attention Module (PWAM)

    model/method

    The Pixel-Word Attention Module (PWAM) densely aligns spatial visual features with linguistic word representations at each stage ii of the vision encoder, generating multi-modal feature maps Fi∈RCi×Hi×WiF_i \in \mathbb{R}^{C_i \times H_i \times W_i} from visual features Vi∈RCi×Hi×WiV_i \in \mathbb{R}^{C_i \times H_i \times W_i} and linguistic features L∈RCt×TL \in \mathbb{R}^{C_t \times T}.

    PWAM operates in two phases:

    1. Pixel-Word Attention Aggregation: Spatial visual features serve as queries to attend over word tokens as keys and values: Viq=flatten(ωiq(Vi))V_{iq} = \text{flatten}(\omega_{iq}(V_i)) Lik=ωik(L)L_{ik} = \omega_{ik}(L) Liv=ωiv(L)L_{iv} = \omega_{iv}(L) Gi′=softmax(ViqTLikCi)LivTG'_i = \text{softmax}\left(\frac{V_{iq}^T L_{ik}}{\sqrt{C_i}}\right) L_{iv}^T Gi=ωiw(unflatten(Gi′T))G_i = \omega_{iw}(\text{unflatten}(G_i^{'T})) where ωik\omega_{ik} and ωiv\omega_{iv} are 1×11 \times 1 convolutions projecting to CiC_i channels, while ωiq\omega_{iq} and ωiw\omega_{iw} are 1×11 \times 1 convolutions followed by instance normalization projecting to CiC_i channels. The operation flatten(⋅)\text{flatten}(\cdot) reshapes (Ci,Hi,Wi)(C_i, H_i, W_i) to (Ci,HiWi)(C_i, H_i W_i) in row-major order, and unflatten(⋅)\text{unflatten}(\cdot) reverses this.

    2. Visual-Linguistic Modality Modulation: Vim=ωim(Vi)V_{im} = \omega_{im}(V_i) Fi=ωio(Vim⊙Gi)F_i = \omega_{io}(V_{im} \odot G_i) where ⊙\odot denotes element-wise multiplication, and ωim\omega_{im} and ωio\omega_{io} are each parameterized as a 1×11 \times 1 convolution followed by a ReLU activation.

    By avoiding cross-attention between two Hi×WiH_i \times W_i spatial feature maps, PWAM maintains a significantly smaller memory footprint than standard spatial-to-spatial cross-attention mechanisms.

  3. Knowl 3 — Language Pathway (LP) and Language Gate (LG)

    model/method

    The Language Pathway (LP) regulates the integration of multi-modal features Fi∈RCi×Hi×WiF_i \in \mathbb{R}^{C_i \times H_i \times W_i} (output by PWAM) back into visual feature maps Vi∈RCi×Hi×WiV_i \in \mathbb{R}^{C_i \times H_i \times W_i} before passing them to the subsequent Transformer encoder stage.

    To prevent multi-modal features from overwhelming pre-trained visual representations, a Language Gate (LG) computes an element-wise spatial scaling map: Si=γi(Fi)S_i = \gamma_i(F_i) Ei=Si⊙Fi+ViE_i = S_i \odot F_i + V_i where ⊙\odot is element-wise multiplication, Ei∈RCi×Hi×WiE_i \in \mathbb{R}^{C_i \times H_i \times W_i} is the input to the next vision Transformer stage ϕi+1\phi_{i+1}, and γi\gamma_i is a two-layer perceptron consisting of a 1×11 \times 1 convolution, a ReLU nonlinearity, a second 1×11 \times 1 convolution, and a hyperbolic tangent (tanh⁡\tanh) activation function.

    Treating language-infused features as additive gated residuals preserves the initialization weights of the vision Transformer pre-trained on unimodal visual data (ImageNet-22K).

  4. Knowl 4 — LAVT Top-Down Segmentation Decoder

    model/method

    LAVT employs a lightweight top-down feature aggregation network to combine multi-scale multi-modal feature maps Fi∈RCi×Hi×WiF_i \in \mathbb{R}^{C_i \times H_i \times W_i} (i∈{1,2,3,4}i \in \{1, 2, 3, 4\}) generated by the stage-wise PWAM modules into final pixel-wise segmentation masks.

    The decoding process is defined recursively from top stage to bottom stage: {Y4=F4,Yi=ρi([υ(Yi+1);Fi]),i=3,2,1\begin{cases} Y_4 = F_4, & \\ Y_i = \rho_i([\upsilon(Y_{i+1}); F_i]), & i = 3, 2, 1 \end{cases} where [⋅;⋅][\cdot ; \cdot] denotes concatenation along the channel dimension, υ(⋅)\upsilon(\cdot) denotes spatial upsampling via bilinear interpolation, and ρi(⋅)\rho_i(\cdot) is a projection module composed of two consecutive 3×33 \times 3 convolutions each followed by batch normalization and ReLU nonlinearity.

    The final fused feature map Y1Y_1 is projected to a 2-channel class score map (target referent vs. background) via a 1×11 \times 1 convolution, with predictions obtained by taking argmax\text{argmax} across the channel dimension.

  5. Knowl 5 — Referring Image Segmentation Performance across RefCOCO, RefCOCO+, and G-Ref

    data/table

    LAVT was evaluated on the validation, testA, and testB splits of RefCOCO and RefCOCO+, and on the UMD and Google splits of G-Ref against prior state-of-the-art referring image segmentation methods. Performance is measured in overall Intersection-over-Union (oIoU, %).

    Method Language
    Model
    RefCOCO RefCOCO+ G-Ref (UMD) G-Ref (Google)
    val test A test B val test A test B val test val
    DMN SRU 49.78 54.83 45.13 38.88 44.22 32.29 - - 36.76
    RRN LSTM 55.33 57.26 53.93 39.75 42.15 36.11 - - 36.45
    MAttNet Bi-LSTM 56.51 62.37 51.70 46.67 52.39 40.08 47.64 48.61 -
    CMSA None 58.32 60.61 55.09 43.76 47.60 37.89 - - 39.98
    CAC Bi-LSTM 58.90 61.77 53.81 - - - 46.37 46.95 44.32
    STEP Bi-LSTM 60.04 63.46 57.97 48.19 52.33 40.41 - - 46.40
    BRINet LSTM 60.98 62.99 59.21 48.17 52.32 42.11 - - 48.04
    CMPC LSTM 61.36 64.53 59.64 49.56 53.44 43.23 - - 49.05
    LSCM LSTM 61.47 64.99 59.55 49.34 53.12 43.50 - - 48.05
    CMPC+ LSTM 62.47 65.08 60.82 50.25 54.04 43.47 - - 49.89
    MCN Bi-GRU 62.44 64.20 59.71 50.62 54.99 44.69 49.22 49.40 -
    EFN Bi-GRU 62.76 65.69 59.67 51.50 55.24 43.01 - - 51.93
    BUSNet Self-Att 63.27 66.41 61.39 51.76 56.87 44.13 - - 50.56
    CGAN Bi-GRU 64.86 68.04 62.07 51.03 55.51 44.06 51.01 51.69 46.54
    LTS Bi-GRU 65.43 67.76 63.08 54.21 58.32 48.02 54.40 54.25 -
    VLT Bi-GRU 65.65 68.29 62.73 55.50 59.20 49.36 52.99 56.65 49.76
    LAVT (Ours) BERT 72.73 75.82 68.79 62.14 68.38 55.10 61.24 62.09 60.50

    LAVT outperforms previous state-of-the-art methods across all splits on all three benchmark datasets without requiring extra multi-task training or separate proposal stages. On the validation sets of RefCOCO, RefCOCO+, G-Ref (UMD), and G-Ref (Google), LAVT establishes absolute improvements of 7.08%, 6.64%, 6.84%, and 8.57% in overall IoU over previous state of the art.

  6. Knowl 6 — Fair Controlled Comparison with Identical Swin-B and BERT Backbones

    data/table

    To isolate the architectural contribution of early encoder-based cross-modal fusion from backbone capacity advantages, LTS, EFN, and VLT were re-implemented and evaluated using the exact same backbones (Swin-B for vision, BERTBASE\text{BERT}_{\text{BASE}} for language) and training recipes as LAVT on the RefCOCO validation set.

    Method [email protected] [email protected] [email protected] oIoU mIoU
    LTS (Swin-B + BERT) 80.59 69.48 26.13 69.94 70.56
    EFN (Swin-B + BERT) 82.55 73.27 31.68 70.76 72.95
    VLT (Swin-B + BERT) 83.24 72.81 24.64 70.89 71.98
    Ours + VLT Decoder 84.57 75.14 26.36 72.12 73.57
    LAVT (Ours) 84.46 75.28 34.30 72.73 74.46

    LAVT outperforms LTS, EFN, and VLT under identical backbones by 1.84% to 2.79% in oIoU and up to 9.66% in high-precision threshold metric [email protected]. Furthermore, replacing LAVT's lightweight decoder with VLT's complex cross-modal Transformer decoder ("Ours + VLT Decoder") degrades oIoU from 72.73% to 72.12% and mIoU from 74.46% to 73.57%, confirming that post-encoding cross-modal decoders are redundant when multi-modal features are effectively fused inside the vision encoder.

  7. Knowl 7 — Ablation Analysis of Core Architectural Components in LAVT

    data/table

    Ablation experiments conducted on the RefCOCO validation set demonstrate the individual impact of the Language Pathway (LP), Pixel-Word Attention Module (PWAM), Language Gate (LG) activation function, PWAM normalization, prediction feature choices, and attention mechanisms.

    Configuration [email protected] [email protected] [email protected] oIoU mIoU
    Full Model (LAVT) 84.46 75.28 34.30 72.73 74.46
    w/o Language Pathway (LP) 81.46 70.80 30.95 70.78 71.96
    w/o PWAM (global sentence pool) 81.76 72.76 32.46 71.03 72.31
    w/o LP and w/o PWAM 77.87 66.93 27.95 68.82 68.87
    LG Activation Function:
    Tanh (Default) 84.46 75.28 34.30 72.73 74.46
    Sigmoid 81.89 72.71 33.35 70.49 72.47
    PWAM Normalization Layer:
    InstanceNorm (Default) 84.46 75.28 34.30 72.73 74.46
    LayerNorm 82.97 74.15 33.99 71.92 73.32
    BatchNorm 82.89 73.82 33.53 71.59 73.09
    None 81.91 72.73 33.11 70.66 72.34
    Attention Mechanism Replacement:
    PWAM (Default) 84.46 75.28 34.30 72.73 74.46
    GA (GARAN) 83.22 74.09 32.71 71.20 73.16
    BCAM 82.26 72.81 33.31 70.19 72.42

    Key takeaways include:

    • Removing the Language Pathway causes a 1.95% drop in oIoU and a 2.50% drop in mIoU.
    • Replacing PWAM with global sentence-level pooling leads to a 1.70% drop in oIoU.
    • Using tanh⁡\tanh rather than Sigmoid in the Language Gate improves oIoU by 2.24%.
    • Instance Normalization inside PWAM outperforms LayerNorm (+0.81% oIoU), BatchNorm (+1.14% oIoU), and no normalization (+2.07% oIoU).
    • PWAM outperforms alternative modules (GA by +1.53% oIoU and BCAM by +2.54% oIoU) with lower computation.

Coverage note — None was omitted; all key architectural components, mathematical formulations, experimental benchmark results, controlled comparative baselines, and comprehensive ablations have been fully represented.

References

  1. 1.Anurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun, Mario Lućić, and Cordelia Schmid. Vivit: A video vision transformer. In ICCV, 2021. 2
  2. 2.Miriam Bellver, Carles Ventura, Carina Silberer, Ioannis Kazakos, Jordi Torres, and Xavier Giro-i Nieto. Refvos: A closer look at referring expressions for video object segmentation. arXiv:2010.00263, 2020. 2
  3. 3.Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. End-to-end object detection with transformers. In ECCV, 2020. 2
  4. 4.Ding-Jie Chen, Songhao Jia, Yi-Chen Lo, Hwann-Tzong Chen, and Tyng-Luh Liu. See-through-text grouping for referring image segmentation. In ICCV, 2019. 1, 2, 6
  5. 5.Jianbo Chen, Yelong Shen, Jianfeng Gao, Jingjing Liu, and Xiaodong Liu. Language-based image editing with recurrent attentive models. In CVPR, 2018. 1
  6. 6.Liang-Chieh Chen, George Papandreou, Florian Schroff, and Hartwig Adam. Rethinking atrous convolution for semantic image segmentation. arXiv:1706.05587, 2017. 2
  7. 7.Yi-Wen Chen, Yi-Hsuan Tsai, Tiantian Wang, Yen-Yu Lin, and Ming-Hsuan Yang. Referring expression object segmentation with caption-aware consistency. In BMVC, 2019. 6
  8. 8.Ming-Ming Cheng, Shuai Zheng, Wen-Yan Lin, Vibhav Vineet, Paul Sturgess, Nigel Crook, Niloy J. Mitra, and Philip Torr. Imagespirit: Verbal guided image parsing. In TOG, 2014. 1
  9. 9.Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc V. Le, and Ruslan Salakhutdinov. Transformer-XL: Attentive language models beyond a fixed-length context. In ACL, 2019. 2
  10. 10.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. ImageNet: A Large-Scale Hierarchical Image Database. In CVPR, 2009. 5
  11. 11.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT: pre-training of deep bidirectional transformers for language understanding. In NAACL, 2019. 2
  12. 12.Henghui Ding, Chang Liu, Suchen Wang, and Xudong Jiang. Vision-language transformer and query generation for referring segmentation. In ICCV, 2021. 1, 2, 6, 8
  13. 13.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition at scale. In ICLR, 2021. 2
  14. 14.Guang Feng, Zhiwei Hu, Lihe Zhang, and Huchuan Lu. Encoder fusion network with co-attention embedding for referring image segmentation. In CVPR, 2021. 2, 4, 6, 8
  15. 15.Raia Hadsell, Sumit Chopra, and Yann LeCun. Dimensionality reduction by learning an invariant mapping. In CVPR, 2006. 2
  16. 16.Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick. Momentum contrast for unsupervised visual representation learning. In CVPR, 2020. 2
  17. 17.Sepp Hochreiter and Jürgen Schmidhuber. Long short-term memory. Neural computation, 1997. 2
  18. 18.Ronghang Hu, Marcus Rohrbach, and Trevor Darrell. Segmentation from natural language expressions. In ECCV, 2016. 1, 2, 4
  19. 19.Ronghang Hu and Amanpreet Singh. Unit: Multimodal multitask learning with a unified transformer. In ICCV, 2021. 2
  20. 20.Zhiwei Hu, Guang Feng, Jiayu Sun, Lihe Zhang, and Huchuan Lu. Bi-directional relationship inferring network for referring image segmentation. In CVPR, 2020. 1, 2, 4, 6, 7, 8
  21. 21.Shaofei Huang, Tianrui Hui, Si Liu, Guanbin Li, Yunchao Wei, Jizhong Han, Luoqi Liu, and Bo Li. Referring image segmentation via cross-modal progressive comprehension. In CVPR, 2020. 2, 6
  22. 22.Tianrui Hui, Si Liu, Shaofei Huang, Guanbin Li, Sansi Yu, Faxi Zhang, and Jizhong Han. Linguistic structure guided context modeling for referring image segmentation. In ECCV, 2020. 2
  23. 23.Tianrui Hui, Si Liu, Shaofei Huang, Guanbin Li, Sansi Yu, Faxi Zhang, and Jizhong Han. Linguistic structure guided context modeling for referring image segmentation. In ECCV, 2020. 6
  24. 24.Sergey Ioffe and Christian Szegedy. Batch normalization: Accelerating deep network training by reducing internal covariate shift. In ICML, 2015. 5
  25. 25.Ya Jing, Tao Kong, Wei Wang, Liang Wang, Lei Li, and Tieniu Tan. Locate then segment: A strong pipeline for referring image segmentation. In CVPR, 2021. 2, 6, 8
  26. 26.Aishwarya Kamath, Mannat Singh, Yann LeCun, Gabriel Synnaeve, Ishan Misra, and Nicolas Carion. Mdetrmodulated detection for end-to-end multi-modal understanding. In ICCV, pages 1780–1790, 2021. 2
  27. 27.Jie Lei, Linjie Li, Luowei Zhou, Zhe Gan, Tamara L. Berg, Mohit Bansal, and Jingjing Liu. Less is more: Clipbert for video-and-language learning via sparse sampling. In CVPR, 2021. 2, 3
  28. 28.Ruiyu Li, Kai-Can Li, Yi-Chun Kuo, Michelle Shu, Xiaojuan Qi, Xiaoyong Shen, and Jiaya Jia. Referring image segmentation via recurrent refinement networks. In CVPR, 2018. 1, 2, 4
  29. 29.Ruiyu Li, Kaican Li, Yi-Chun Kuo, Michelle Shu, Xiaojuan Qi, Xiaoyong Shen, and Jiaya Jia. Referring image segmentation via recurrent refinement networks. In CVPR, 2018. 6
  30. 30.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C. Lawrence Zitnick. Microsoft coco: Common objects in context. In David Fleet, Tomas Pajdla, Bernt Schiele, and Tinne Tuytelaars, editors, ECCV, 2014. 1, 5
  31. 31.Chenxi Liu, Zhe Lin, Xiaohui Shen, Jimei Yang, Xin Lu, and Alan Yuille. Recurrent multimodal interaction for referring image segmentation. In ICCV, 2017. 1, 2
  32. 32.Si Liu, Tianrui Hui, Shaofei Huang, Yunchao Wei, Bo Li, and Guanbin Li. Cross-modal progressive comprehension for referring segmentation. In TPAMI, 2021. 4, 6
  33. 33.Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. In ICCV, 2021. 2, 3, 5
  34. 34.Ze Liu, Jia Ning, Yue Cao, Yixuan Wei, Zheng Zhang, Stephen Lin, and Han Hu. Video swin transformer. arXiv:2106.13230, 2021. 2
  35. 35.Jonathan Long, Evan Shelhamer, and Trevor Darrell. Fully convolutional networks for semantic segmentation. In CVPR, 2015. 2
  36. 36.Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. In ICLR, 2019. 5
  37. 37.Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks. In NeurIPS, 2019. 2, 3
  38. 38.Gen Luo, Yiyi Zhou, Rongrong Ji, Xiaoshuai Sun, Jinsong Su, Chia-Wen Lin, and Qi Tian. Cascade grouped attention network for referring expression segmentation. In ACMMM, 2020. 4, 6, 7, 8
  39. 39.Gen Luo, Yiyi Zhou, Xiaoshuai Sun, Liujuan Cao, Chenglin Wu, Cheng Deng, and Rongrong Ji. Multi-task collaborative network for joint referring expression comprehension and segmentation. In CVPR, 2020. 2, 6, 7, 8
  40. 40.Junhua Mao, Jonathan Huang, Alexander Toshev, Oana Camburu, Alan L Yuille, and Kevin Murphy. Generation and comprehension of unambiguous object descriptions. In CVPR, 2016. 2, 5, 6
  41. 41.Edgar Margffoy-Tuay, Juan C Perez, Emilio Botero, and Pablo Arbeláez. Dynamic multimodal instance segmentation guided by natural language queries. In ECCV, 2018. 4, 6
  42. 42.Varun K. Nagaraja, Vlad I. Morariu, and Larry S. Davis. Modeling context between objects for referring expression understanding. In ECCV, 2016. 2, 5, 6
  43. 43.Vinod Nair and Geoffrey E. Hinton. Rectified linear units improve restricted boltzmann machines. In ICML, 2010. 4, 5
  44. 44.Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al. Pytorch: An imperative style, high-performance deep learning library. In NeurIPS, 2019. 5
  45. 45.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning transferable visual models from natural language supervision. In ICML, 2021. 2
  46. 46.Yongming Rao, Wenliang Zhao, Guangyi Chen, Yansong Tang, Zheng Zhu, Guan Huang, Jie Zhou, and Jiwen Lu. Denseclip: Language-guided dense prediction with contextaware prompting. In CVPR, 2022. 2
  47. 47.Joseph Redmon and Ali Farhadi. Yolov3: An incremental improvement. arXiv:1804.02767, 2018. 2
  48. 48.Hengcan Shi, Hongliang Li, Fanman Meng, and Qingbo Wu. Key-word-aware network for referring expression image segmentation. In ECCV, 2018. 1, 2
  49. 49.Hengcan Shi, Hongliang Li, Fanman Meng, and Qingbo Wu. Key-word-aware network for referring expression image segmentation. In ECCV, 2018. 4
  50. 50.Kihyuk Sohn. Improved deep metric learning with multi-class n-pair loss objective. In NeurIPS, 2016. 2
  51. 51.Robin Strudel, Ricardo Garcia, Ivan Laptev, and Cordelia Schmid. Segmenter: Transformer for semantic segmentation. In ICCV, 2021. 2
  52. 52.Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Herve Jegou. Training data-efficient image transformers & distillation through attention. In ICML, 2021. 2
  53. 53.Dmitry Ulyanov, Andrea Vedaldi, and Victor Lempitsky. Instance normalization: The missing ingredient for fast stylization. arXiv:1607.08022, 2016. 4
  54. 54.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. In NeurIPS, 2017. 2, 4, 5
  55. 55.Xin Wang, Qiuyuan Huang, Asli Celikyilmaz, Jianfeng Gao, Dinghan Shen, Yuan-Fang Wang, William Yang Wang, and Lei Zhang. Reinforced cross-modal matching and selfsupervised imitation learning for vision-language navigation. In CVPR, 2019. 1
  56. 56.Thomas Wolf, Julien Chaumond, Lysandre Debut, Victor Sanh, Clement Delangue, Anthony Moi, Pierric Cistac, Morgan Funtowicz, Joe Davison, Sam Shleifer, et al. Transformers: State-of-the-art natural language processing. In EMNLP, 2020. 5
  57. 57.Sibei Yang, Meng Xia, Guanbin Li, Hong-Yu Zhou, and Yizhou Yu. Bottom-up shift and reasoning for referring image segmentation. In CVPR, 2021. 6
  58. 58.Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Ruslan Salakhutdinov, and Quoc V. Le. XLNet: Generalized autoregressive pretraining for language understanding. In NeurIPS, 2019. 2
  59. 59.Zhao Yang, Yansong Tang, Luca Bertinetto, Hengshuang Zhao, and Philip H.S. Torr. Hierarchical interaction network for video object segmentation from referring expressions. In BMVC, 2021. 6
  60. 60.Linwei Ye, Mrigank Rochan, Zhi Liu, and Yang Wang. Cross-modal self-attention network for referring image segmentation. In CVPR, 2019. 2, 4
  61. 61.Linwei Ye, Mrigank Rochan, Zhi Liu, and Yang Wang. Cross-modal self-attention network for referring image segmentation. In CVPR, 2019. 6
  62. 62.Licheng Yu, Zhe Lin, Xiaohui Shen, Jimei Yang, Xin Lu, Mohit Bansal, and Tamara L Berg. Mattnet: Modular attention network for referring expression comprehension. In CVPR, 2018. 2, 6
  63. 63.Licheng Yu, Patrick Poirson, Shan Yang, Alexander C Berg, and Tamara L Berg. Modeling context in referring expressions. In ECCV, 2016. 2, 5, 6
  64. 64.Sixiao Zheng, Jiachen Lu, Hengshuang Zhao, Xiatian Zhu, Zekun Luo, Yabiao Wang, Yanwei Fu, Jianfeng Feng, Tao Xiang, Philip HS Torr, et al. Rethinking semantic segmentation from a sequence-to-sequence perspective with transformers. In CVPR, 2021. 2
  65. 65.Bolei Zhou, Hang Zhao, Xavier Puig, Tete Xiao, Sanja Fidler, Adela Barriuso, and Antonio Torralba. Semantic understanding of scenes through the ade20k dataset. IJCV, 2019. 1
  66. 66.Xizhou Zhu, Weijie Su, Lewei Lu, Bin Li, Xiaogang Wang, and Jifeng Dai. Deformable DETR: deformable transformers for end-to-end object detection. In ICLR, 2021. 2

Citation

MLA
Yang, Z., et al. “LAVT: Language-Aware Vision Transformer for Referring Image Segmentation”. arXiv, 2021, http://arxiv.org/abs/2112.02244v2.
APA
Yang, Z., Wang, J., Tang, Y., Chen, K., Zhao, H., & Torr, P. H. S. (2021). LAVT: Language-Aware Vision Transformer for Referring Image Segmentation. arXiv. http://arxiv.org/abs/2112.02244v2
Chicago
Yang, Z., J. Wang, Y. Tang, K. Chen, H. Zhao, and P. H. S. Torr. 2021. “LAVT: Language-Aware Vision Transformer for Referring Image Segmentation”. arXiv. http://arxiv.org/abs/2112.02244v2.
Harvard
Yang, Z. et al. (2021) “LAVT: Language-Aware Vision Transformer for Referring Image Segmentation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2112.02244v2.
Vancouver
1. Yang Z, Wang J, Tang Y, Chen K, Zhao H, Torr PHS (2021) LAVT: Language-Aware Vision Transformer for Referring Image Segmentation. arXiv

BibTeX

@article{yang2021lavt,
  title = {LAVT: Language-Aware Vision Transformer for Referring Image Segmentation},
  author = {Yang, Zhao and Wang, Jiaqi and Tang, Yansong and Chen, Kai and Zhao, Hengshuang and Torr, Philip H. S.},
  year = {2021},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2112.02244v2},
  eprint = {2112.02244}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE