ReSTR: Convolution-free Referring Image Segmentation Using Transformers
Namyup KimDongwon KimSuha KwakCuiling LanWenjun Zeng
Proposes ReSTR, the first purely transformer-based, convolution-free model for referring image segmentation that unifies vision and language processing through self-attention to capture long-range cross-modal dependencies and achieve state-of-the-art performance across major benchmarks.
Real-world computer vision systems increasingly require the ability to identify and segment specific image regions based on free-form natural language queries rather than predefined categories. This capability is vital for interactive photo editing, robotics, and assistive technologies. However, traditional systems rely heavily on convolutional and recurrent neural networks, which inherently struggle to process long-range contextual relationships within sentences and fail to flexibly fuse visual and textual information.
The article demonstrates the design and performance of ReSTR, the first entirely convolution-free framework for referring image segmentation that relies exclusively on attention-based transformer architectures. The primary objective is to evaluate whether a unified transformer network can improve cross-modal comprehension, handle complex linguistic relationships, and deliver superior segmentation accuracy.
The researchers evaluated ReSTR against existing state-of-the-art models across four standard benchmark datasets: ReferIt, UNC, UNC+, and Gref. The architecture processes non-overlapping image patches and word embeddings through separate transformer encoders, joins them using a specialized indirect multimodal fusion encoder that prevents visual bias, and generates final pixel-level segmentations through a lightweight coarse-to-fine decoding module.
The evaluation yielded several key findings. First, ReSTR established top-tier performance across public benchmarks, achieving intersection-over-union scores of 70.18% on ReferIt, 67.22% on UNC validation, 55.78% on UNC+ validation, and 54.48% on Gref validation. Second, the model demonstrated exceptional resilience when processing long, complex language expressions: on the Gref dataset, performance dropped by only 6.81 percentage points between short and long phrases, compared to a 13.71-point drop in competitive prior methods. Third, computational efficiency was substantially improved, requiring only 52.29 billion multiply-accumulate operations (MACs)—less than half the computational load of leading alternatives—while simultaneously eliminating the need for slow post-processing techniques.
These results confirm that removing convolutional constraints allows visual and textual features to interact with greater flexibility and precision. Organizations developing vision-language applications can achieve higher accuracy with lower compute overhead, reducing deployment latency and operational infrastructure costs without requiring complex auxiliary pipelines.
Based on these findings, development teams should consider adopting unified transformer-based pipelines for multimodal segmentation tasks. For production environments with strict memory constraints, adopting weight sharing within the fusion module is recommended, as it halves parameter count with negligible impact on accuracy. Future development should focus on integrating linear-complexity transformer architectures to mitigate the quadratic compute scaling associated with smaller visual patch sizes, which remains the primary computational bottleneck of the model.
- Paper: Generation and Comprehension of Unambiguous Object Descriptions, Junhua Mao et al. (2015). This seminal work establishes the foundational dataset and formulation for referring object comprehension that ReSTR directly aims to solve with transformers.
- Paper: Rethinking Semantic Segmentation from a Sequence-to-Sequence Perspective with Transformers, Sixiao Zheng et al. (2020). This study introduces the sequence-to-sequence transformer paradigm for pixel-level visual segmentation, laying the architectural groundwork for ReSTR's convolution-free design.
- Paper: Segmenter: Transformer for Semantic Segmentation, Robin Strudel et al. (2021). This paper establishes the pure Vision Transformer encoder-decoder framework for dense segmentation, which informs ReSTR's vision processing strategy.
- Paper: UNITER: UNiversal Image-TExt Representation Learning, Yen-Chun Chen et al. (2020). This foundational paper presents universal cross-modal transformer fusion for vision-and-language tasks, motivating ReSTR's self-attention multi-modal interaction mechanism.
- Paper: SegFormer: Simple and Efficient Design for Semantic Segmentation with Transformers, Enze Xie et al. (2021). This paper introduces lightweight transformer designs and multi-scale feature encoding for dense image segmentation, which contextually underpins convolution-free vision modeling.
- Paper: Per-Pixel Classification is Not All You Need for Semantic Segmentation, Bowen Cheng et al. (2021). This work unifies segmentation through transformer decoders and mask prediction, providing fundamental concepts utilized in modern attention-based segmentation architectures.
- Paper: Twins: Revisiting the Design of Spatial Attention in Vision Transformers, Xiangxiang Chu et al. (2021). This work explores spatial attention in vision transformer backbones to handle dense prediction tasks efficiently.
- Paper: Flickr30k Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence Models, Bryan A. Plummer et al. (2015). This dataset and benchmark paper introduces region-to-phrase grounding essential for learning natural language visual references.
- Paper: LAVT: Language-Aware Vision Transformer for Referring Image Segmentation, Zhao Yang et al. (2022). Published concurrently, this paper advances transformer-based referring image segmentation by fusing linguistic representations directly inside the vision transformer encoder stages.
- Paper: GSVA: Generalized Segmentation via Multimodal Large Language Models, Zhuofan Xia et al. (2024). This work generalizes referring segmentation beyond single-target constraints to multi-target and non-existent entity scenarios by leveraging multimodal large language models.
- Paper: End-to-End Referring Video Object Segmentation with Multimodal Transformers, Adam Botach et al. (2022). This paper extends multimodal transformer-based referring segmentation architectures from static images into temporal video sequences.
- Paper: Learning Open-Vocabulary Semantic Segmentation Models From Natural Language Supervision, Jilan Xu et al. (2023). This work expands language-supervised visual segmentation into open-vocabulary settings learned purely from web image-text pairs.
- Paper: CLIP as RNN: Segment Countless Visual Concepts without Training Endeavor, Shuyang Sun et al. (2024). This study advances language-driven referring and semantic segmentation to a training-free, recurrent vision-language paradigm.
