ReSTR: Convolution-free Referring Image Segmentation Using Transformers

Namyup KimDongwon KimSuha KwakCuiling LanWenjun Zeng

article2022CVPR210 citations

Proposes ReSTR, the first purely transformer-based, convolution-free model for referring image segmentation that unifies vision and language processing through self-attention to capture long-range cross-modal dependencies and achieve state-of-the-art performance across major benchmarks.

Listen

Real-world computer vision systems increasingly require the ability to identify and segment specific image regions based on free-form natural language queries rather than predefined categories. This capability is vital for interactive photo editing, robotics, and assistive technologies. However, traditional systems rely heavily on convolutional and recurrent neural networks, which inherently struggle to process long-range contextual relationships within sentences and fail to flexibly fuse visual and textual information.

The article demonstrates the design and performance of ReSTR, the first entirely convolution-free framework for referring image segmentation that relies exclusively on attention-based transformer architectures. The primary objective is to evaluate whether a unified transformer network can improve cross-modal comprehension, handle complex linguistic relationships, and deliver superior segmentation accuracy.

The researchers evaluated ReSTR against existing state-of-the-art models across four standard benchmark datasets: ReferIt, UNC, UNC+, and Gref. The architecture processes non-overlapping image patches and word embeddings through separate transformer encoders, joins them using a specialized indirect multimodal fusion encoder that prevents visual bias, and generates final pixel-level segmentations through a lightweight coarse-to-fine decoding module.

The evaluation yielded several key findings. First, ReSTR established top-tier performance across public benchmarks, achieving intersection-over-union scores of 70.18% on ReferIt, 67.22% on UNC validation, 55.78% on UNC+ validation, and 54.48% on Gref validation. Second, the model demonstrated exceptional resilience when processing long, complex language expressions: on the Gref dataset, performance dropped by only 6.81 percentage points between short and long phrases, compared to a 13.71-point drop in competitive prior methods. Third, computational efficiency was substantially improved, requiring only 52.29 billion multiply-accumulate operations (MACs)—less than half the computational load of leading alternatives—while simultaneously eliminating the need for slow post-processing techniques.

These results confirm that removing convolutional constraints allows visual and textual features to interact with greater flexibility and precision. Organizations developing vision-language applications can achieve higher accuracy with lower compute overhead, reducing deployment latency and operational infrastructure costs without requiring complex auxiliary pipelines.

Based on these findings, development teams should consider adopting unified transformer-based pipelines for multimodal segmentation tasks. For production environments with strict memory constraints, adopting weight sharing within the fusion module is recommended, as it halves parameter count with negligible impact on accuracy. Future development should focus on integrating linear-complexity transformer architectures to mitigate the quadratic compute scaling associated with smaller visual patch sizes, which remains the primary computational bottleneck of the model.

arXiv: 2203.16768
Cover for ReSTR: Convolution-free Referring Image Segmentation Using Transformers

Abstract

Referring image segmentation is an advanced semantic segmentation task where target is not a predefined class but is described in natural language. Most of existing methods for this task rely heavily on convolutional neural networks, which however have trouble capturing long-range dependencies between entities in the language expression and are not flexible enough for modeling interactions between the two different modalities. To address these issues, we present the first convolution-free model for referring image segmentation using transformers, dubbed ReSTR. Since it extracts features of both modalities through transformer encoders, it can capture long-range dependencies between entities within each modality. Also, ReSTR fuses features of the two modalities by a self-attention encoder, which enables flexible and adaptive interactions between the two modalities in the fusion process. The fused features are fed to a segmentation module, which works adaptively according to the image and language expression in hand. ReSTR is evaluated and compared with previous work on all public benchmarks, where it outperforms all existing models.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 2.1. Semantic Segmentation
  • 2.2. Referring Image Segmentation
  • 2.3. Vision Transformer
  • 3. Proposed Method
  • 3.1. Visual and Linguistic Feature Extraction
  • 3.2. Multimodal Fusion Encoder
  • 3.3. Coarse-to-Fine Segmentation Decoder
  • 4. Experiments
  • 4.1. Experimental Setting
  • 4.2. Comparisons with the State of the Art
  • 4.3. Analysis of Variants of Fusion Encoder
  • 4.4. In-depth Analysis of ReSTR
  • 5. Conclusion
  • References

Knowls

  1. Knowl 1 — ReSTR convolution-free referring image segmentation architecture

    model/method

    ReSTR is a convolution-free referring image segmentation network that receives an image and a natural-language expression and predicts the pixels referred to by the expression. The image is represented as non-overlapping patch embeddings and the expression as word embeddings. ReSTR first applies separate transformer encoders to the visual and linguistic sequences, then uses a transformer-based multimodal fusion encoder to produce patch-wise multimodal features and an expression-dependent adaptive classifier. A coarse-to-fine decoder applies the adaptive classifier to image patches and progressively upsamples the resulting representation to a pixel-level mask. The architecture therefore models long-range dependencies within each modality, flexible interactions between modalities, and language-conditioned segmentation without convolutional or recurrent layers.

  2. Knowl 2 — Transformer-based visual and linguistic feature extraction

    model/method

    For an image xv∈RH×W×Cvx^v\in\mathbb{R}^{H\times W\times C_v}, ReSTR splits the image into non-overlapping square patches of side length PP, yielding Nv=HW/P2N_v=HW/P^2 patch tokens with projected channel dimension DvD_v. A learnable positional encoding Eposv∈RNv×DvE^v_{\mathrm{pos}}\in\mathbb{R}^{N_v\times D_v} is added before a vision transformer produces visual features zv=Transformers⁡(xp+Eposv;θv)z_v=\operatorname{Transformers}(x_p+E^v_{\mathrm{pos}};\theta_v). A language expression is represented by NlN_l word embeddings xl∈RNl×Clx_l\in\mathbb{R}^{N_l\times C_l}, where NlN_l is the padded maximum sentence length; a sinusoidal positional encoding is added before a language transformer produces linguistic features zlz_l.

    Each transformer encoder contains MM sequential transformer blocks. For an input sequence zi∈RN×Dz_i\in\mathbb{R}^{N\times D} at block ii, the block is

    zˉi+1=MSA⁡(LN⁡(zi))+zi,\bar z_{i+1}=\operatorname{MSA}(\operatorname{LN}(z_i))+z_i, zi+1=MLP⁡(LN⁡(zˉi+1))+zˉi+1.z_{i+1}=\operatorname{MLP}(\operatorname{LN}(\bar z_{i+1}))+\bar z_{i+1}.

    Here NN is the number of tokens, DD is the channel dimension, LN⁡\operatorname{LN} is layer normalization, and MSA is multi-head self-attention. For attention head hh with query, key, and value matrices qh,kh,vh∈RN×Dhq_h,k_h,v_h\in\mathbb{R}^{N\times D_h}, the attention operation is

    SA⁡h(z)=Ahvh,Ah=softmax⁡(qhkhTDh),\operatorname{SA}_h(z)=A_hv_h,\qquad A_h=\operatorname{softmax}\left(\frac{q_hk_h^{\mathsf T}}{\sqrt{D_h}}\right),

    and MSA⁡(z)=[SA⁡1(z),…,SA⁡k(z)]WMSA\operatorname{MSA}(z)=[\operatorname{SA}_1(z),\ldots,\operatorname{SA}_k(z)]W_{\mathrm{MSA}}, where kk is the number of heads, Dh=D/kD_h=D/k, and WMSA∈RkDh×DW_{\mathrm{MSA}}\in\mathbb{R}^{kD_h\times D} is a learned output projection. Self-attention gives both modality-specific encoders access to global interactions from the beginning of feature extraction.

  3. Knowl 3 — Indirect conjugating multimodal fusion and adaptive classifier

    model/method

    ReSTR aligns the visual and linguistic features to a common channel dimension DD and fuses them with two transformer encoders. The visual-linguistic encoder processes the visual and linguistic sequences in parallel:

    [zv′,zl′]=Transformers⁡([zv,zl];θvl),[z'_v,z'_l]=\operatorname{Transformers}([z_v,z_l];\theta_{vl}),

    where zv′∈RNv×Dz'_v\in\mathbb{R}^{N_v\times D} is the patch-wise multimodal representation and zl′∈RNl×Dz'_l\in\mathbb{R}^{N_l\times D} is the visual-attended linguistic representation. A trainable class seed embedding es∈R1×De_s\in\mathbb{R}^{1\times D} is then processed together with the visual-attended linguistic features by a second transformer encoder:

    es′=Transformers⁡([zl′,es];θls), e'_s=\operatorname{Transformers}([z'_l,e_s];\theta_{ls}),

    where es′∈R1×De'_s\in\mathbb{R}^{1\times D} is an adaptive classifier for the entity described by the expression. This indirect interaction makes the linguistic features the medium connecting the class seed to visual information: the classifier can use image-dependent appearance information carried by the visual-attended language features while avoiding direct attention from the seed to irrelevant visual patches. This design is called the indirect Conjugating Multimodal Encoder (CME).

  4. Knowl 4 — Coarse-to-fine segmentation decoder and training objective

    algorithm

    ReSTR first produces a patch-level probability for each of the NvN_v image patches by comparing each patch representation with the expression-dependent classifier:

    y^p=σ(zv′(es′)TD),\hat y_p=\sigma\left(\frac{z'_v(e'_s)^{\mathsf T}}{\sqrt D}\right),

    where zv′∈RNv×Dz'_v\in\mathbb{R}^{N_v\times D}, es′∈R1×De'_s\in\mathbb{R}^{1\times D}, y^p∈[0,1]Nv×1\hat y_p\in[0,1]^{N_v\times1}, and σ\sigma is the element-wise sigmoid. The patch probabilities mask the multimodal features by broadcast Hadamard multiplication, zmasked=zv′⊙y^pz_{\mathrm{masked}}=z'_v\odot\hat y_p. The decoder concatenates the masked multimodal features with the projected visual features and applies K=log⁡2PK=\log_2P sequential blocks. Each block upsamples the spatial token grid by a factor of two, linearly reduces the channel dimension by one half, and applies an activation function. A final linear projection reshapes the output into the pixel-level prediction Y^m∈[0,1]H×W×1\hat Y_m\in[0,1]^{H\times W\times1}; only this pixel-level prediction is used at inference.

    For training, the ground-truth mask Ym∈{0,1}H×W×1Y_m\in\{0,1\}^{H\times W\times1} is average-pooled within each patch pip_i and thresholded at τ\tau to create a patch label ypiy_p^i:

    ypi={1,h(pi)>τ,0,otherwise, y_p^i=\begin{cases}1,&h(p_i)>\tau,\\0,&\text{otherwise},\end{cases}

    where h(pi)h(p_i) is the spatial average of the ground-truth pixels in patch pip_i. The total loss combines binary cross-entropy on the patch prediction and the pixel prediction:

    L=λ Lb(y^p,yp)+Lb(Y^m,Ym).\mathcal{L}=\lambda\,\mathcal{L}_b(\hat y_p,y_p)+\mathcal{L}_b(\hat Y_m,Y_m).

    The experiments use τ=0.8\tau=0.8 and λ=0.1\lambda=0.1. The decoder is intended to refine the coarse localization supplied by patch classification into a high-resolution mask.

  5. Knowl 5 — Datasets, implementation, and evaluation protocol

    experimental setup

    ReSTR is evaluated on four referring-image-segmentation benchmarks. ReferIt contains 19,894 images, 130,525 expressions, and 96,654 masks. UNC contains 19,994 images, 142,209 expressions, and 50,000 masks. UNC+ contains 19,992 images, 141,564 expressions, and 49,856 masks; unlike UNC, its expressions exclude location words and use appearance descriptions only. Gref contains 25,711 images, 104,560 expressions, and 54,822 objects.

    The vision encoder is an ImageNet-21K-pretrained ViT-B/16 with 12 transformer layers, patch size 16, 768 channels, 12 attention heads, and 3,072-dimensional MLP expansion. The language encoder uses pretrained 300-dimensional GloVe word embeddings, six transformer layers, 12 attention heads, and 3,072-dimensional MLP expansion; expression length is padded or truncated to Nl=20N_l=20. The multimodal fusion encoder uses the same transformer configuration as the vision encoder, and the segmentation decoder has four blocks because P=16P=16. Input images are resized to 480×480480\times480.

    All models are trained with AdamW, weight decay 5×10−45\times10^{-4}, initial learning rate 10−510^{-5} with polynomial decay, batch size 8, and 400,000 iterations, including a 40,000-iteration warm-up. The evaluation metric is cumulative intersection-over-union, computed as total intersection divided by total union over all test samples. Accuracy is additionally reported at IoU thresholds 0.50.5, 0.60.6, 0.70.7, 0.80.8, and 0.90.9.

  6. Knowl 6 — State-of-the-art benchmark performance

    data/table

    The benchmark comparison measures cumulative IoU (%) on the four referring-segmentation datasets and their official validation or test splits. ReSTR uses no DenseCRF post-processing and obtains the best result on every reported split except UNC+ testB, where VLT is higher. It improves over prior methods while using an ImageNet-classification-pretrained visual backbone rather than a COCO-detection-pretrained backbone.

    Method DCRF ReferIt test UNC val UNC testA UNC testB UNC+ val UNC+ testA UNC+ testB Gref val
    LSTM-CNN – 48.03 – – – – – – 28.14
    RMI ✓ 58.73 45.18 45.69 45.57 29.86 30.48 29.50 34.52
    DMN – 52.81 49.78 54.83 45.13 38.88 44.22 32.29 36.76
    RRN ✓ 63.63 55.33 57.26 53.95 39.75 42.15 36.11 36.45
    CMSA ✓ 63.80 58.32 60.61 55.09 43.76 47.60 37.89 39.98
    STEP – 64.13 60.04 63.46 57.97 48.19 52.33 40.41 46.40
    BRINet ✓ 63.46 61.35 63.37 59.57 48.57 52.87 42.13 48.04
    LSCM ✓ 66.57 61.47 64.99 59.55 49.34 53.12 43.50 48.05
    CMPC ✓ 65.53 61.36 64.54 59.64 49.56 53.44 43.23 49.05
    ACM – 66.70 62.76 65.69 59.67 51.50 55.24 43.01 51.93
    BUSNet ✓ – 63.27 66.41 61.39 51.76 56.87 44.13 50.56
    LTS – – 65.43 67.76 63.08 54.21 58.32 48.02 54.40
    VLT – – 65.65 68.29 62.73 55.50 59.20 49.36 52.99
    ReSTR – 70.18 67.22 69.30 64.45 55.78 60.44 48.27 54.48

    ReSTR reaches 70.18 on ReferIt test, 67.22/69.30/64.45 on UNC val/testA/testB, 55.78/60.44/48.27 on UNC+ val/testA/testB, and 54.48 on Gref val.

  7. Knowl 7 — Robustness to expression length

    empirical result

    ReSTR retains a larger fraction of its segmentation accuracy for long referring expressions than the comparison methods, supporting the use of transformer self-attention for long-range cross-modal interactions. On Gref, ReSTR obtains IoUs of 58.72, 53.47, 53.96, and 51.91 for expression lengths 1–5, 6–7, 8–10, and 11–20, respectively. On the same split, ACM falls from 59.92 for lengths 1–5 to 46.21 for lengths 11–20, a decrease of 13.71 percentage points, whereas ReSTR decreases by 6.81 points.

    ReSTR’s IoUs across the remaining length partitions are 72.38, 69.46, 61.19, and 50.21 on UNC for lengths 1–2, 3, 4–5, and 6–20; 65.72, 54.81, 47.65, and 37.02 on UNC+ for lengths 1–2, 3, 4–5, and 6–20; and 80.82, 69.78, 63.66, and 50.73 on ReferIt for lengths 1, 2, 3–4, and 5–20. ReSTR is better than the compared methods for most length groups, with the exception of the shortest 1–5 group on Gref, where ACM reaches 59.92 versus ReSTR’s 58.72.

  8. Knowl 8 — Fusion-encoder design ablation

    data/table

    The multimodal fusion design was evaluated on the Gref validation set using four transformer layers. A Vanilla Multimodal Encoder (VME) feeds visual features, linguistic features, and the class seed into one transformer simultaneously. Because the visual sequence has 900 tokens while the language sequence has 20 tokens in the experiment, the class seed’s attention is strongly biased toward visual features: the visual/language attention scores are 82.4%/16.8% in layer 1, 98.9%/1.0% in layer 2, 98.7%/1.2% in layer 3, and 98.1%/1.7% in layer 4. An Independent Multimodal Encoder (IME) prevents direct interaction between the class seed and visual features, whereas CME connects them indirectly through visual-attended linguistic features. CME gives the best IoU, and weight sharing nearly halves the fusion-encoder parameter count with only a small accuracy decrease.

    Fusion encoder Parameters MACs IoU (%)
    VME 28.35M 31.36G 51.27
    IME 28.35M 15.96G 45.89
    CME 28.35M 15.96G 52.81
    CME with weight sharing 14.18M 15.96G 52.79

    The results show that merely allowing all sequences to attend jointly is inferior to the indirect linguistic mediation used by CME, while direct removal of seed–visual interaction is also ineffective.

  9. Knowl 9 — Effects of fusion depth, decoder refinement, and weight sharing

    data/table

    An ablation on Gref validation evaluates the number of transformer layers in the multimodal fusion encoder, the coarse-to-fine decoder, and weight sharing. The fusion encoder contains two transformer encoders, so its total layer count is tested at 2, 4, or 6. Four layers provide the strongest non-shared result. Adding the decoder improves IoU from 52.81 to 54.48 with four fusion layers, but only from 48.12 to 48.43 with two layers, indicating that the decoder is most useful when the patch-level representation is sufficiently accurate. Sharing transformer weights causes only a small degradation: 54.48 to 54.07 with four layers and 52.84 to 52.59 with six layers.

    Fusion layers Weight sharing Decoder [email protected] [email protected] [email protected] [email protected] [email protected] IoU (%)
    2 – – 52.60 45.59 36.59 23.54 5.23 48.12
    2 – ✓ 52.86 46.61 38.93 26.37 7.90 48.43
    4 – – 61.77 55.86 46.86 30.88 8.18 52.81
    4 – ✓ 64.91 59.94 51.73 37.70 12.23 54.48
    4 ✓ ✓ 64.27 59.01 50.70 35.85 11.46 54.07
    6 – – 63.36 57.88 48.75 33.46 8.75 52.84
    6 ✓ – 63.05 57.32 48.19 32.47 8.47 52.59

    The decoder analysis with six fusion layers was not performed because of memory limitations.

  10. Knowl 10 — Computational efficiency and patch-size limitation

    limitation

    On Gref validation, ReSTR achieves higher IoU with substantially lower computation than recent referring-segmentation systems. MACs in this comparison are measured with a 320×320320\times320 input, and the DenseCRF column indicates whether a post-processing step is used.

    Method DCRF Parameters MACs IoU (%)
    BRINet ✓ 241.18M 367.63G 48.04
    LSCM ✓ 127.91M 130.45G 48.05
    CMPC ✓ 118.66M 126.66G 49.05
    ACM 232.78M 124.68G 51.93
    ReSTR (CME) 122.87M 52.29G 54.48
    ReSTR (CME with weight sharing) 108.70M 52.29G 54.07

    The principal limitation is the quadratic dependence of transformer computation on the number of image patches. Decreasing patch size increases the token count and therefore sharply increases computational cost, while dense-prediction quality depends strongly on patch size. ReSTR consequently has an accuracy–efficiency trade-off, and the paper identifies linear-complexity transformer architectures as a future way to alleviate it.

Coverage note — Qualitative mask visualizations were omitted because they illustrate the already quantified coarse-to-fine behavior rather than add a separate methodological or numerical result.

References

  1. 1.Gedas Bertasius, Lorenzo Torresani, Stella X. Yu, and Jianbo Shi. Convolutional random walk networks for semantic image segmentation. In Proc. IEEE Conference on Computer Vision and Pattern Recognition (CVPR), July 2017.
  2. 2.Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. End-to-end object detection with transformers. In Proc. European Conference on Computer Vision (ECCV), 2020.
  3. 3.Ding-Jie Chen, Songhao Jia, Yi-Chen Lo, Hwann-Tzong Chen, and Tyng-Luh Liu. See-through-text grouping for referring image segmentation. In Proc. IEEE International Conference on Computer Vision (ICCV), 2019.
  4. 4.Liang-Chieh Chen, George Papandreou, Iasonas Kokkinos, Kevin Murphy, and Alan L Yuille. Semantic image segmentation with deep convolutional nets and fully connected CRFs. In Proc. International Conference on Learning Representations (ICLR), 2015.
  5. 5.Liang-Chieh Chen, George Papandreou, Iasonas Kokkinos, Kevin Murphy, and Alan L. Yuille. Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs. IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), 2017.
  6. 6.Liang-Chieh Chen, Yukun Zhu, George Papandreou, Florian Schroff, and Hartwig Adam. Encoder-decoder with atrous separable convolution for semantic image segmentation. In Proc. European Conference on Computer Vision (ECCV), 2018.
  7. 7.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. ImageNet: a large-scale hierarchical image database. In Proc. IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2009.
  8. 8.Henghui Ding, Chang Liu, Suchen Wang, and Xudong Jiang. Vision-language transformer and query generation for referring segmentation. In Proc. IEEE International Conference on Computer Vision (ICCV), 2021.
  9. 9.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. In Proc. International Conference on Learning Representations (ICLR), 2021.
  10. 10.Hugo Jair Escalante, Carlos A Hernandez, Jesus A Gonzalez, Aurelio Lopez-López, Manuel Montes, Eduardo F Morales, L Enrique Sucar, Luis Villasenor, and Michael Grubinger. The segmented and annotated iapr tc-12 benchmark. Computer vision and image understanding, 2010.
  11. 11.Guang Feng, Zhiwei Hu, Lihe Zhang, and Huchuan Lu. Encoder fusion network with co-attention embedding for referring image segmentation. In Proc. IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2021.
  12. 12.Ronghang Hu, Marcus Rohrbach, and Trevor Darrell. Segmentation from natural language expressions. In Proc. European Conference on Computer Vision (ECCV), 2016.
  13. 13.Zhiwei Hu, Guang Feng, Jiayu Sun, Lihe Zhang, and Huchuan Lu. Bi-directional relationship inferring network for referring image segmentation. In Proc. IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2020.
  14. 14.Shaofei Huang, Tianrui Hui, Si Liu, Guanbin Li, Yunchao Wei, Jizhong Han, Luoqi Liu, and Bo Li. Referring image segmentation via cross-modal progressive comprehension. In Proc. IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2020.
  15. 15.Zilong Huang, Xinggang Wang, Lichao Huang, Chang Huang, Yunchao Wei, and Wenyu Liu. Ccnet: Criss-cross attention for semantic segmentation. In Proc. IEEE International Conference on Computer Vision (ICCV), 2019.
  16. 16.Tianrui Hui, Si Liu, Shaofei Huang, Guanbin Li, Sansi Yu, Faxi Zhang, and Jizhong Han. Linguistic structure guided context modeling for referring image segmentation. In Proc. European Conference on Computer Vision (ECCV), 2020.
  17. 17.Andrew Jaegle, Felix Gimeno, Andrew Brock, Andrew Zisserman, Oriol Vinyals, and Joao Carreira. Perceiver: General perception with iterative attention. In Proc. International Conference on Machine Learning (ICML), 2021.
  18. 18.Ya Jing, Tao Kong, Wei Wang, Liang Wang, Lei Li, and Tieniu Tan. Locate then segment: A strong pipeline for referring image segmentation. In Proc. IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2021.
  19. 19.Aishwarya Kamath, Mannat Singh, Yann LeCun, Gabriel Synnaeve, Ishan Misra, and Nicolas Carion. Mdetr-modulated detection for end-to-end multi-modal understanding. In Proc. IEEE International Conference on Computer Vision (ICCV), 2021.
  20. 20.Sahar Kazemzadeh, Vicente Ordonez, Mark Matten, and Tamara Berg. Referitgame: Referring to objects in photographs of natural scenes. In Proc. Empirical Methods in Natural Language Processing (EMNLP), 2014.
  21. 21.Philipp Krahenbühl and Vladlen Koltun. Efficient inference in fully connected crfs with gaussian edge potentials. In Proc. Neural Information Processing Systems (NeurIPS), 2011.
  22. 22.Ruiyu Li, Kaican Li, Yi-Chun Kuo, Michelle Shu, Xiaojuan Qi, Xiaoyong Shen, and Jiaya Jia. Referring image segmentation via recurrent refinement networks. In Proc. IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018.
  23. 23.Guosheng Lin, Anton Milan, Chunhua Shen, and Ian Reid. Refinenet: Multi-path refinement networks for high-resolution semantic segmentation. In Proc. IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017.
  24. 24.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollar, and C Lawrence Zitnick. Microsoft COCO: common objects in context. In Proc. European Conference on Computer Vision (ECCV), 2014.
  25. 25.Chenxi Liu, Zhe Lin, Xiaohui Shen, Jimei Yang, Xin Lu, and Alan Yuille. Recurrent multimodal interaction for referring image segmentation. In Proc. IEEE International Conference on Computer Vision (ICCV), 2017.
  26. 26.Z. Liu, X. Li, P. Luo, C. C. Loy, and X. Tang. Semantic image segmentation via deep parsing network. In Proc. IEEE International Conference on Computer Vision (ICCV), 2015.
  27. 27.Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. In Proc. IEEE International Conference on Computer Vision (ICCV), 2021.
  28. 28.Jonathan Long, Evan Shelhamer, and Trevor Darrell. Fully convolutional networks for semantic segmentation. In Proc. IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2015.
  29. 29.Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. Proc. International Conference on Learning Representations (ICLR), 2019.
  30. 30.Junhua Mao, Jonathan Huang, Alexander Toshev, Oana Camburu, Alan L Yuille, and Kevin Murphy. Generation and comprehension of unambiguous object descriptions. In Proc. IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016.
  31. 31.Edgar Margffoy-Tuay, Juan C Perez, Emilio Botero, and Pablo Arbelaez. Dynamic multimodal instance segmentation guided by natural language queries. In Proc. European Conference on Computer Vision (ECCV), 2018.
  32. 32.Arsha Nagrani, Shan Yang, Anurag Arnab, Aren Jansen, Cordelia Schmid, and Chen Sun. Attention bottlenecks for multimodal fusion. In Proc. Neural Information Processing Systems (NeurIPS), 2021.
  33. 33.Hyeonwoo Noh, Seunghoon Hong, and Bohyung Han. Learning deconvolution network for semantic segmentation. In Proc. IEEE International Conference on Computer Vision (ICCV), 2015.
  34. 34.Jeffrey Pennington, Richard Socher, and Christopher D Manning. Glove: Global vectors for word representation. In Proc. Empirical Methods in Natural Language Processing (EMNLP), 2014.
  35. 35.Prajit Ramachandran, Niki Parmar, Ashish Vaswani, Irwan Bello, Anselm Levskaya, and Jonathon Shlens. Stand-alone self-attention in vision models. Proc. Neural Information Processing Systems (NeurIPS), 2019.
  36. 36.Olaf Ronneberger, Philipp Fischer, and Thomas Brox. U-net: Convolutional networks for biomedical image segmentation. In Proc. Medical Image Computing and Computer-Assisted Intervention (MICCAI), 2015.
  37. 37.Hengcan Shi, Hongliang Li, Fanman Meng, and Qingbo Wu. Key-word-aware network for referring expression image segmentation. In Proc. European Conference on Computer Vision (ECCV), 2018.
  38. 38.Aravind Srinivas, Tsung-Yi Lin, Niki Parmar, Jonathon Shlens, Pieter Abbeel, and Ashish Vaswani. Bottleneck transformers for visual recognition. In Proc. IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2021.
  39. 39.Robin Strudel, Ricardo Garcia, Ivan Laptev, and Cordelia Schmid. Segmenter: Transformer for semantic segmentation. In Proc. IEEE International Conference on Computer Vision (ICCV), 2021.
  40. 40.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Proc. Neural Information Processing Systems (NeurIPS), 2017.
  41. 41.Huiyu Wang, Yukun Zhu, Hartwig Adam, Alan Yuille, and Liang-Chieh Chen. Max-deeplab: End-to-end panoptic segmentation with mask transformers. In Proc. IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2021.
  42. 42.Wenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan, Kaitao Song, Ding Liang, Tong Lu, Ping Luo, and Ling Shao. Pyramid vision transformer: A versatile backbone for dense prediction without convolutions. Proc. IEEE International Conference on Computer Vision (ICCV), 2021.
  43. 43.Xiaolong Wang, Ross Girshick, Abhinav Gupta, and Kaiming He. Non-local neural networks. In Proc. IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018.
  44. 44.Maoke Yang, Kun Yu, Chi Zhang, Zhiwei Li, and Kuiyuan Yang. Denseaspp for semantic segmentation in street scenes. In Proc. IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018.
  45. 45.Sibei Yang, Meng Xia, Guanbin Li, Hong-Yu Zhou, and Yizhou Yu. Bottom-up shift and reasoning for referring image segmentation. In Proc. IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2021.
  46. 46.Linwei Ye, Mrigank Rochan, Zhi Liu, and Yang Wang. Cross-modal self-attention network for referring image segmentation. In Proc. IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2019.
  47. 47.Fisher Yu and Vladlen Koltun. Multi-scale context aggregation by dilated convolutions. In Proc. International Conference on Learning Representations (ICLR), 2016.
  48. 48.Licheng Yu, Patrick Poirson, Shan Yang, Alexander C Berg, and Tamara L Berg. Modeling context in referring expressions. In Proc. European Conference on Computer Vision (ECCV), 2016.
  49. 49.Matthew D Zeiler, Graham W Taylor, and Rob Fergus. Adaptive deconvolutional networks for mid and high level feature learning. In Proc. IEEE International Conference on Computer Vision (ICCV), 2011.
  50. 50.Fan Zhang, Yanqin Chen, Zhihang Li, Zhibin Hong, Jingtuo Liu, Feifei Ma, Junyu Han, and Errui Ding. Acfnet: Attentional class feature network for semantic segmentation. In Proc. IEEE International Conference on Computer Vision (ICCV), pages 6798–6807, 2019.
  51. 51.Hengshuang Zhao, Jiaya Jia, and Vladlen Koltun. Exploring self-attention for image recognition. In Proc. IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2020.
  52. 52.H. Zhao, J. Shi, X. Qi, X. Wang, and J. Jia. Pyramid scene parsing network. In Proc. IEEE Conference on Computer Vision and Pattern Recognition (CVPR), July 2017.
  53. 53.Hengshuang Zhao, Yi Zhang, Shu Liu, Jianping Shi, Chen Change Loy, Dahua Lin, and Jiaya Jia. Psanet: Point-wise spatial attention network for scene parsing. In Proc. European Conference on Computer Vision (ECCV), 2018.
  54. 54.Shuai Zheng, Sadeep Jayasumana, Bernardino Romera-Paredes, Vibhav Vineet, Zhizhong Su, Dalong Du, Chang Huang, and Philip Torr. Conditional random fields as recurrent neural networks. In Proc. IEEE International Conference on Computer Vision (ICCV), 2015.
  55. 55.Sixiao Zheng, Jiachen Lu, Hengshuang Zhao, Xiatian Zhu, Zekun Luo, Yabiao Wang, Yanwei Fu, Jianfeng Feng, Tao Xiang, Philip HS Torr, et al. Rethinking semantic segmentation from a sequence-to-sequence perspective with transformers. In Proc. IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2021.

Citation

MLA
Kim, N., et al. “ReSTR: Convolution-free Referring Image Segmentation Using Transformers”. arXiv, 2022, http://arxiv.org/abs/2203.16768v1.
APA
Kim, N., Kim, D., Lan, C., Zeng, W., & Kwak, S. (2022). ReSTR: Convolution-free Referring Image Segmentation Using Transformers. arXiv. http://arxiv.org/abs/2203.16768v1
Chicago
Kim, N., D. Kim, C. Lan, W. Zeng, and S. Kwak. 2022. “ReSTR: Convolution-free Referring Image Segmentation Using Transformers”. arXiv. http://arxiv.org/abs/2203.16768v1.
Harvard
Kim, N. et al. (2022) “ReSTR: Convolution-free Referring Image Segmentation Using Transformers”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2203.16768v1.
Vancouver
1. Kim N, Kim D, Lan C, Zeng W, Kwak S (2022) ReSTR: Convolution-free Referring Image Segmentation Using Transformers. arXiv

BibTeX

@article{kim2022restr,
  title = {ReSTR: Convolution-free Referring Image Segmentation Using Transformers},
  author = {Kim, Namyup and Kim, Dongwon and Lan, Cuiling and Zeng, Wenjun and Kwak, Suha},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2203.16768v1},
  eprint = {2203.16768}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE