End-to-End Referring Video Object Segmentation with Multimodal Transformers

Adam BotachEvgenii ZheltonozhskiiChaim Baskin

article2022CVPR242 citations

Proposes Multimodal Tracking Transformer (MTTR), an end-to-end framework that models referring video object segmentation as parallel sequence prediction to achieve state-of-the-art accuracy at 76 frames per second without relying on complex post-processing or specialized inductive biases.

Listen

Identifying and tracking a specific target in video using a natural language description is a core requirement for next-generation automated video analysis. However, standard methods in this domain have historically relied on complicated, multi-stage pipelines that separate language comprehension, visual object detection, tracking, and boundary refinement into distinct modules. These complex setups struggle with action-based language descriptions, suffer when objects are temporarily obscured, and introduce operational inefficiencies that hinder real-time deployment.

The article demonstrates an end-to-end artificial intelligence framework that significantly simplifies this task. The objective is to establish whether a unified architecture based on an attention-driven model can process natural language queries and video frames simultaneously to accurately identify, track, and segment referenced targets without auxiliary post-processing or specialized language heuristics.

To achieve this, the authors developed the Multimodal Tracking Transformer. The system extracts visual features from video frames and linguistic features from text queries, projecting both into a shared sequence that is processed by a single multimodal network. Instead of locating only the referenced entity, the model tracks all candidate objects in parallel across video frames and generates high-resolution segmentation masks. It then applies a temporal segment voting mechanism to score each tracked object sequence based on how strongly it corresponds to the textual description across the entire video. The approach was evaluated against standard benchmark datasets, including A2D-Sentences, JHMDB-Sentences, and the public validation benchmark of Refer-YouTube-VOS.

The evaluation revealed substantial performance and efficiency improvements over existing state-of-the-art approaches. First, on the primary benchmark, the model achieved a 5.7-point gain in mean Average Precision and a 6.7% absolute improvement in overlap accuracy compared to previous leading methods. Second, the system operated at an inference speed of 76 frames per second on a single standard graphics processing unit, proving its capability for real-time processing. Third, when tested directly on an un-finetuned dataset to evaluate generalization, it outperformed prior systems by 5.0 points in mean Average Precision. Fourth, on the larger and more complex benchmark, it attained leading accuracy metrics even though it was trained on less data and operated without model ensembles or task-specific pre-training.

These findings indicate that removing complex multi-stage pipelines in favor of a unified sequence-prediction framework lowers architectural complexity while raising segmentation accuracy and inference speed. For operational systems, this approach reduces computational overhead and minimizes integration risks associated with maintaining separate detection, tracking, and mask-refinement components. It also demonstrates that standard cross-entropy objectives and temporal voting sufficiently align text with video actions without complicated inductive modules.

For practical implementation, engineering teams should evaluate this unified sequence-prediction design when developing automated video retrieval and tracking workflows. Organizations seeking to optimize performance should consider temporal window configurations around ten frames, which provided optimal accuracy in benchmark testing. Further engineering work should involve validating the model on domain-specific edge cases, such as video streams with long-term target occlusions or dense crowds, before deploying to production environments.

Confidence in these findings is high across the evaluated benchmark conditions due to consistent gains across multiple metrics and datasets. However, decision-makers should note certain limitations: performance drops slightly when temporal context windows become too large or when relying on non-contextual word embeddings, and low-precision edge annotations in older benchmark datasets can constrain fine-grained evaluation.

Cover for End-to-End Referring Video Object Segmentation with Multimodal Transformers

Abstract

The referring video object segmentation task (RVOS) involves segmentation of a text-referred object instance in the frames of a given video. Due to the complex nature of this multimodal task, which combines text reasoning, video understanding, instance segmentation and tracking, existing approaches typically rely on sophisticated pipelines in order to tackle it. In this paper, we propose a simple Transformer-based approach to RVOS. Our framework, termed Multimodal Tracking Transformer (MTTR), models the RVOS task as a sequence prediction problem. Following recent advancements in computer vision and natural language processing, MTTR is based on the realization that video and text can be processed together effectively and elegantly by a single multimodal Transformer model. MTTR is end-to-end trainable, free of text-related inductive bias components and requires no additional mask-refinement post-processing steps. As such, it simplifies the RVOS pipeline considerably compared to existing methods. Evaluation on standard benchmarks reveals that MTTR significantly outperforms previous art across multiple metrics. In particular, MTTR shows impressive +5.7 and +5.0 mAP gains on the A2D-Sentences and JHMDB-Sentences datasets respectively, while processing 76 frames per second. In addition, we report strong results on the public validation set of Refer-YouTube-VOS, a more challenging RVOS dataset that has yet to receive the attention of researchers. The code to reproduce our experiments is available at https://github.com/mttr2021/MTTR.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Method
  • 3.1. Method Overview
  • 3.2. Temporal Encoder
  • 3.3. Multimodal Transformer
  • 3.4. The Instance Segmentation Process
  • 3.5. Instance Sequence Matching
  • 3.6. Loss Functions
  • 3.7. Inference
  • 4. Experiments
  • 4.1. Implementation Details
  • 4.2. Comparison with State-of-the-Art Methods
  • 4.3. Ablation Studies
  • 4.4. Qualitative Analysis
  • 5. Conclusion
  • References

Knowls

  1. Knowl 1 — Multimodal Tracking Transformer (MTTR) Architecture for Referring Video Object Segmentation

    model/method

    The Multimodal Tracking Transformer (MTTR) models Referring Video Object Segmentation (RVOS) as an end-to-end parallel sequence prediction problem. Given an input video frame sequence V={vi}i=1TV = \{v_i\}_{i=1}^T with vi∈RC×H0×W0v_i \in \mathbb{R}^{C \times H_0 \times W_0} and a linguistic query T={ti}i=1LT = \{t_i\}_{i=1}^L, the task requires predicting segmentation masks for the referred entity over a subset of interest frames VI⊆VV_I \subseteq V of length TIT_I.

    MTTR executes this process through three stages:

    1. Feature Extraction: Visual features fVIt∈RH×W×CVf_{V_I}^t \in \mathbb{R}^{H \times W \times C_V} are extracted per frame using a Video Swin Transformer temporal backbone (retaining the first 3 blocks and modifying temporal downsampling to output per-frame representations). Linguistic features fT∈RL×DTf_T \in \mathbb{R}^{L \times D_T} are extracted using a Transformer text encoder (such as RoBERTa-base). Both features are linearly projected to a shared embedding dimension DD.
    2. Multimodal Sequence Generation and Transformer Processing: For each frame of interest, flattened visual tokens are concatenated with the projected text tokens to create TIT_I multimodal sequences of shape (H⋅W+L)×D(H \cdot W + L) \times D. These sequences are processed in parallel by a Transformer encoder to model cross-modal spatio-temporal dependencies. A Transformer decoder takes NqN_q learned object query sequences (where queries across different frames sharing the same parameter index attend to the same instance across time, enabling natural tracking) and attends to the multimodal features to produce decoded instance features Q={qt}t=1TI∈RTI×DQ = \{q_t\}_{t=1}^{T_I} \in \mathbb{R}^{T_I \times D} for each sequence.
    3. Instance Segmentation and Reference Prediction: For each instance query sequence, a dynamic segmentation kernel head generates conditional convolution kernels to generate spatio-temporal mask sequences, while a reference classification head predicts whether the sequence corresponds to the text-referred object.
  2. Knowl 2 — Dynamic Kernel Mask Generation with FPN Spatial Decoder in MTTR

    model/method

    MTTR generates instance mask sequences without requiring auxiliary mask refinement post-processing (such as DenseCRFs). The segmentation process operates as follows:

    1. Let FEF_E denote the output of the final Transformer encoder layer. The visual tokens corresponding to the frames of interest are extracted and reshaped into feature maps FEVI∈RTI×H×W×DF_E^{V_I} \in \mathbb{R}^{T_I \times H \times W \times D}.
    2. The visual features FEVIF_E^{V_I} are hierarchically fused with the multi-scale intermediate feature maps FB1,…,n−1F_B^{1,\dots,n-1} from the first n−1n-1 blocks of the temporal encoder using an FPN-like spatial decoder GSegG_{Seg} to construct high-resolution feature maps: FSeg={fSegt}t=1TI,fSegt∈RDs×H04×W04F_{Seg} = \{f_{Seg}^t\}_{t=1}^{T_I}, \quad f_{Seg}^t \in \mathbb{R}^{D_s \times \frac{H_0}{4} \times \frac{W_0}{4}}
    3. For each decoded instance query sequence Q={qt}t=1TIQ = \{q_t\}_{t=1}^{T_I} with qt∈RDq_t \in \mathbb{R}^D, a two-layer perceptron GkernelG_{kernel} produces a sequence of conditional segmentation kernels: Gkernel(Q)={kt}t=1TI,kt∈RDsG_{kernel}(Q) = \{k_t\}_{t=1}^{T_I}, \quad k_t \in \mathbb{R}^{D_s}
    4. The predicted mask sequence M={mt}t=1TIM = \{m_t\}_{t=1}^{T_I} is obtained by convolving each dynamic kernel with its corresponding frame feature map, followed by bilinear interpolation to the original image dimensions H0×W0H_0 \times W_0: mt=Upsample(kt∗fSegt)∈RH0×W0m_t = \text{Upsample}(k_t * f_{Seg}^t) \in \mathbb{R}^{H_0 \times W_0}
  3. Knowl 3 — Bipartite Sequence Matching for Video Instance Association in MTTR

    model/method

    To train MTTR with sequence-level supervision, a bipartite matching problem is solved between the set of NqN_q predicted instance sequences y^={y^j}j=1Nq\hat{y} = \{\hat{y}_j\}_{j=1}^{N_q} and the ground-truth sequences y={yi}i=1Nqy = \{y_i\}_{i=1}^{N_q} (padded with ∅\emptyset if the number of ground-truth instances Ni<NqN_i < N_q).

    Each ground-truth sequence is represented as: yi=(mi,ri)=({mit}t=1TI,{rit}t=1TI)y_i = (m_i, r_i) = \left( \{m_i^t\}_{t=1}^{T_I}, \{r_i^t\}_{t=1}^{T_I} \right) where mit∈{0,1}H0×W0m_i^t \in \{0, 1\}^{H_0 \times W_0} is the ground-truth mask at frame tt, and rit∈{0,1}2r_i^t \in \{0, 1\}^2 is a one-hot vector indicating whether instance ii is the text-referred object and visible in frame tt. If sequence ii is padding, mi=∅m_i = \emptyset.

    Each predicted sequence is defined as: y^j=(m^j,r^j)=({m^jt}t=1TI,{r^jt}t=1TI)\hat{y}_j = (\hat{m}_j, \hat{r}_j) = \left( \{\hat{m}_j^t\}_{t=1}^{T_I}, \{\hat{r}_j^t\}_{t=1}^{T_I} \right) where r^jt=GRef(qjt)∈[0,1]2\hat{r}_j^t = G_{Ref}(q_j^t) \in [0, 1]^2 is output by a linear classification head GRefG_{Ref} followed by a softmax function.

    The pair-wise matching cost between predicted sequence y^j\hat{y}_j and target sequence yiy_i is: CMatch(y^j,yi)=1{mi≠∅}[λdCDice(m^j,mi)+λrCRef(r^j,ri)]\mathcal{C}_{Match}(\hat{y}_j, y_i) = \mathbb{1}_{\{m_i \neq \emptyset\}} \left[ \lambda_d \mathcal{C}_{Dice}(\hat{m}_j, m_i) + \lambda_r \mathcal{C}_{Ref}(\hat{r}_j, r_i) \right] where λd,λr∈R\lambda_d, \lambda_r \in \mathbb{R} are hyperparameters, CDice\mathcal{C}_{Dice} computes the average negative Dice coefficient across all time steps, and CRef\mathcal{C}_{Ref} is defined as: CRef(r^j,ri)=−1TI∑t=1TIr^jt⋅rit\mathcal{C}_{Ref}(\hat{r}_j, r_i) = -\frac{1}{T_I} \sum_{t=1}^{T_I} \hat{r}_j^t \cdot r_i^t

    The optimal assignment permutation σ^∈SNq\hat{\sigma} \in \mathcal{S}_{N_q} is computed using the Hungarian algorithm: σ^=arg⁡min⁡σ∈SNq∑i=1NqCMatch(y^σ(i),yi)\hat{\sigma} = \arg\min_{\sigma \in \mathcal{S}_{N_q}} \sum_{i=1}^{N_q} \mathcal{C}_{Match}(\hat{y}_{\sigma(i)}, y_i)

  4. Knowl 4 — MTTR Optimization Loss Formulation

    equation

    Given the optimal permutation σ^\hat{\sigma} from bipartite sequence matching, the loss function for MTTR over the matched predicted sequences y^\hat{y} and target sequences yy is defined as: L(y^,y)=∑i=1Nq[1{mi≠∅}LMask(m^i,mi)+LRef(r^i,ri)]\mathcal{L}(\hat{y}, y) = \sum_{i=1}^{N_q} \left[ \mathbb{1}_{\{m_i \neq \emptyset\}} \mathcal{L}_{Mask}(\hat{m}_i, m_i) + \mathcal{L}_{Ref}(\hat{r}_i, r_i) \right]

    The mask loss term LMask\mathcal{L}_{Mask} combines the Dice loss and per-pixel Focal loss normalized by the number of instances in the batch: LMask(m^i,mi)=λdLDice(m^i,mi)+λfLFocal(m^i,mi)\mathcal{L}_{Mask}(\hat{m}_i, m_i) = \lambda_d \mathcal{L}_{Dice}(\hat{m}_i, m_i) + \lambda_f \mathcal{L}_{Focal}(\hat{m}_i, m_i) where λd,λf∈R\lambda_d, \lambda_f \in \mathbb{R} are loss scaling hyperparameters.

    The reference classification loss LRef\mathcal{L}_{Ref} supervises the sequence-level alignment between text and object predictions via cross-entropy: LRef(r^i,ri)=−λr1TI∑t=1TIrit⋅log⁡(r^it)\mathcal{L}_{Ref}(\hat{r}_i, r_i) = -\lambda_r \frac{1}{T_I} \sum_{t=1}^{T_I} r_i^t \cdot \log(\hat{r}_i^t) where λr∈R\lambda_r \in \mathbb{R} is a weight hyperparameter matching that of the assignment cost CMatch\mathcal{C}_{Match}. The loss component corresponding to the negative ("un-referred") class is downweighted by a factor of 10 to counteract class imbalance between referred and non-referred background queries.

  5. Knowl 5 — Temporal Segment Voting Scheme (TSVS) for Video-Level Sequence Selection

    algorithm

    During inference, MTTR uses the Temporal Segment Voting Scheme (TSVS) to select the predicted instance sequence that best corresponds to the referring text query by accumulating per-frame class probabilities across the entire temporal sequence.

    Input: Set of predicted reference probability sequences R={r^i}i=1Nq\mathcal{R} = \{\hat{r}_i\}_{i=1}^{N_q}, where each r^i={r^it}t=1TI\hat{r}_i = \{\hat{r}_i^t\}_{t=1}^{T_I}; corresponding predicted segmentation mask sequences M={Mi}i=1Nq\mathcal{M} = \{M_i\}_{i=1}^{N_q}, where Mi={mit}t=1TIM_i = \{m_i^t\}_{t=1}^{T_I}.
    Output: Predicted segmentation mask sequence MpredM_{pred} for the referred instance.
    for each sequence i∈{1,…,Nq}i \in \{1, \dots, N_q\} do
        Si←0S_i \leftarrow 0
        for each frame t∈{1,…,TI}t \in \{1, \dots, T_I\} do
            pref(r^it)←p_{ref}(\hat{r}_i^t) \leftarrow probability of positive ("referred") class from r^it\hat{r}_i^t
            Si←Si+pref(r^it)S_i \leftarrow S_i + p_{ref}(\hat{r}_i^t)
        end for
    end for
    i∗←arg⁡max⁡i∈{1,…,Nq}Sii^* \leftarrow \arg\max_{i \in \{1, \dots, N_q\}} S_i
    Mpred←Mi∗M_{pred} \leftarrow M_{i^*}
    return MpredM_{pred}

    By aggregating probabilities over the full temporal sequence, TSVS allows the model to rely primarily on frames where the referred entity is clearly visible and dismiss ambiguous frames where the object is occluded or outside the field of view.

  6. Knowl 6 — Performance Comparison on A2D-Sentences and JHMDB-Sentences Benchmarks

    data/table

    MTTR was evaluated on the A2D-Sentences and JHMDB-Sentences datasets (evaluating JHMDB-Sentences in a zero-shot transfer setting without fine-tuning). Performance is measured using Precision@KK (K∈[0.5,0.9]K \in [0.5, 0.9]), Overall IoU, Mean IoU, and mean Average Precision (mAP) computed over 0.50:0.05:0.95.

    Method Precision IoU mAP
    50% 60% 70% 80% 90% Overall Mean
    A2D-Sentences Dataset
    Hu et al. 34.8 23.6 13.3 3.3 0.1 47.4 35.0 13.2
    Gavrilyuk et al. (RGB) 47.5 34.7 21.1 8.0 0.2 53.6 42.1 19.8
    RefVOS 57.8 – – – 9.3 67.2 49.7 –
    AAMN 68.1 62.9 52.3 29.6 2.9 61.7 55.2 39.6
    CMSA+CFSA 48.7 43.1 35.8 23.1 5.2 61.8 43.2 –
    CSTM 65.4 58.9 49.7 33.3 9.1 66.2 56.1 39.9
    CMPC-V (I3D) 65.5 59.2 50.6 34.2 9.8 65.3 57.3 40.4
    MTTR (w=8w=8) 72.1 68.4 60.7 45.6 16.4 70.2 61.8 44.7
    MTTR (w=10w=10) 75.4 71.2 63.8 48.5 16.9 72.0 64.0 46.1
    JHMDB-Sentences Dataset
    Hu et al. 63.3 35.0 8.5 0.2 0.0 54.6 52.8 17.8
    Gavrilyuk et al. (RGB) 69.9 46.0 17.3 1.4 0.0 54.1 54.2 23.3
    AAMN 77.3 62.7 36.0 4.4 0.0 58.3 57.6 32.1
    CMSA+CFSA 76.4 62.5 38.9 9.0 0.1 62.8 58.1 –
    CSTM 78.3 63.9 37.8 7.6 0.0 59.8 60.4 33.5
    CMPC-V (I3D) 81.3 65.7 37.1 7.0 0.0 61.6 61.7 34.2
    MTTR (w=8w=8) 91.0 81.5 57.0 14.4 0.1 67.4 67.9 36.6
    MTTR (w=10w=10) 93.9 85.2 61.6 16.6 0.1 70.1 69.8 39.2

    On A2D-Sentences, MTTR (w=10w=10) outperforms previous state-of-the-art CMPC-V by +5.7+5.7 mAP, +6.7%+6.7\% Mean IoU, and +6.7%+6.7\% Overall IoU while operating at 76 frames per second on a single RTX 3090 GPU. On JHMDB-Sentences, MTTR improves upon CMPC-V by +5.0+5.0 mAP and +8.1%+8.1\% Mean IoU.

  7. Knowl 7 — Performance on Refer-YouTube-VOS Benchmark

    data/table

    MTTR was evaluated on the public validation set of Refer-YouTube-VOS containing full-video referring expressions. Evaluation metrics include region similarity (J\mathcal{J}), contour accuracy (F\mathcal{F}), and their average (\mathcal{J}\\&\mathcal{F}).

    Method
    F J\mathcal{J} F\mathcal{F}
    Evaluated on original validation set
    URVOS 47.23 45.27 49.19
    CMPC-V (I3D) 47.48 45.64 49.32
    Evaluated on public validation set
    Ding et al.+^+ 54.80 53.70 56.00
    MTTR (ours) 55.32 54.00 56.64

    Note: +^+ denotes model ensemble and training with external datasets (COCO, RefCOCO, YouTube-VOS).

    Without using external segmentation pretraining on COCO or model ensembling, MTTR outperforms the competitive ensemble approach of Ding et al. on the public validation benchmark, obtaining an overall \mathcal{J}\\&\mathcal{F} score of 55.3255.32.

  8. Knowl 8 — Ablation Studies on Visual Backbones, Temporal Context Window, and Text Encoders

    data/table

    Ablations on A2D-Sentences analyze the effects of 2D vs. 3D visual feature extractors, temporal context window size (ww), and text encoder representations.

    Configuration Overall IoU Mean IoU mAP
    (a) Visual Backbones (without temporal context / static frame)
    CMPC-I (DeepLab-ResNet101) 64.9 51.5 35.1
    MTTR (DeepLab-ResNet101) 67.5 60.2 41.2
    MTTR (Video Swin-T, w=1w=1) 68.9 60.3 41.8
    (b) Input Window Size ww (Video Swin-T + RoBERTa-base)
    w=1w = 1 68.9 60.3 41.8
    w=4w = 4 69.7 61.5 43.8
    w=6w = 6 69.5 61.8 44.0
    w=8w = 8 70.2 61.8 44.7
    w=10w = 10 72.0 64.0 46.1
    w=12w = 12 69.3 62.0 44.0
    (c) Linguistic Embeddings (w=6w = 6)
    RoBERTa (base) 69.5 61.8 44.0
    BERT (base) 69.7 62.1 44.3
    Distil-RoBERTa (base) 70.5 62.4 43.8
    GloVe 68.3 60.9 43.4
    fastText 67.3 59.5 43.1

    Key takeaways:

    1. Even without temporal context (w=1w=1), MTTR with DeepLab-ResNet101 (41.241.2 mAP) surpasses the 2D baseline CMPC-I (35.135.1 mAP) by +6.1+6.1 mAP.
    2. Expanding the temporal context window from w=1w=1 to w=10w=10 provides a +4.3+4.3 mAP gain (41.8→46.141.8 \to 46.1), whereas further expansion to w=12w=12 leads to performance degradation (44.044.0 mAP).
    3. Contextual Transformer text models (RoBERTa, BERT, Distil-RoBERTa) demonstrate consistent performance (43.8–44.343.8\text{--}44.3 mAP), outperforming static word embeddings (GloVe and fastText).
  9. Knowl 9 — Role of Supervising Un-Referred Instances in MTTR Training Stability

    empirical result

    When video samples contain non-referred ground-truth instances alongside the text-referred object, supervising their detection as negative (un-referred) sequences during bipartite matching is essential for stable convergence.

    Ablation experiments omitting supervision of un-referred instances revealed that the model consistently collapses into a degenerate local minimum of the reference loss LRef\mathcal{L}_{Ref}. In this degenerate state, a single object query repeatedly matches all ground-truth instances across samples, leaving the remaining object query slots untrained. While training occasionally escapes this local minimum after several epochs in some configurations, the optimization failure frequently leads to significant degradation in final segmentation mAP.

Coverage note — None was omitted; all key architectural components, loss and matching formulations, inference schemes, empirical benchmarks, and ablation findings from the paper are fully covered.

References

  1. 1.Miriam Bellver, Carles Ventura, Carina Silberer, Jordi Torres Ioannis Kazakos, and Xavier Giró-i-Nieto. RefVOS: a closer look at referring expressions for video object segmentation. arXiv preprint arXiv:2010.00263, 2020. (cited on p. 6)
  2. 2.Piotr Bojanowski, Edouard Grave, Armand Joulin, and Tomas Mikolov. Enriching word vectors with subword information. Transactions of the Association for Computational Linguistics, 5:135–146, 06 2017. (cited on p. 7)
  3. 3.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D. Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. Language models are few-shot learners. In H. Larochelle, M. Ranzato, R. Hadsell, M. F. Balcan, and H. Lin, editors, Advances in Neural Information Processing Systems, volume 33, pages 1877–1901. Curran Associates, Inc., 2020. (cited on pp. 1 and 2)
  4. 4.Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. End-to-end object detection with transformers. In Andrea Vedaldi, Horst Bischof, Thomas Brox, and Jan-Michael Frahm, editors, European Conference on Computer Vision (ECCV), pages 213–229. Springer, August 2020. (cited on pp. 1, 2, 3, 4, and 5)
  5. 5.Joao Carreira and Andrew Zisserman. Quo vadis, action recognition? A new model and the kinetics dataset. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), July 2017. (cited on pp. 2 and 4)
  6. 6.Liang-Chieh Chen, George Papandreou, Iasonas Kokkinos, Kevin Murphy, and Alan L. Yuille. DeepLab: semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected CRFs. IEEE Transactions on Pattern Analysis and Machine Intelligence, 40(4):834–848, 2017. (cited on p. 7)
  7. 7.Jifeng Dai, Haozhi Qi, Yuwen Xiong, Yi Li, Guodong Zhang, Han Hu, and Yichen Wei. Deformable convolutional networks. In IEEE International Conference on Computer Vision (ICCV), Oct 2017. (cited on p. 2)
  8. 8.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT: Pre-training of deep bidirectional transformers for language understanding. In Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota, June 2019. Association for Computational Linguistics. (cited on pp. 1, 2, and 7)
  9. 9.Zihan Ding, Tianrui Hui, Shaofei Huang, Si Liu, Xuan Luo, Junshi Huang, and Xiaoming Wei. Progressive multimodal interaction network for referring video object segmentation. The 3rd Large-scale Video Object Segmentation Challenge, June 2021. (cited on p. 7)
  10. 10.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition at scale. In International Conference on Learning Representations, 2021. (cited on pp. 1, 2, and 3)
  11. 11.Mark Everingham, Luc Van Gool, Christopher K. I. Williams, John Winn, and Andrew Zisserman. The PASCAL visual object classes (VOC) challenge. International Journal of Computer Vision, 88(2):303–338, 2010. (cited on p. 7)
  12. 12.Kirill Gavrilyuk, Amir Ghodrati, Zhenyang Li, and Cees G. M. Snoek. Actor and action video segmentation from a sentence. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 2018. (cited on pp. 2, 4, 5, 6, 7, and 8)
  13. 13.Ronghang Hu, Marcus Rohrbach, and Trevor Darrell. Segmentation from natural language expressions. In Bastian Leibe, Jiri Matas, Nicu Sebe, and Max Welling, editors, European Conference on Computer Vision (ECCV), pages 108–124, Cham, 2016. Springer, Springer International Publishing. (cited on p. 6)
  14. 14.Tianrui Hui, Shaofei Huang, Si Liu, Zihan Ding, Guanbin Li, Wenguan Wang, Jizhong Han, and Fei Wang. Collaborative spatial-temporal modeling for language-queried video actor segmentation. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 4187–4196, June 2021. (cited on pp. 2, 4, 6, and 7)
  15. 15.Hueihan Jhuang, Juergen Gall, Silvia Zuffi, Cordelia Schmid, and Michael J. Black. Towards understanding action recognition. In IEEE International Conference on Computer Vision (ICCV), pages 3192–3199, December 2013. (cited on pp. 5 and 7)
  16. 16.Aishwarya Kamath, Mannat Singh, Yann LeCun, Gabriel Synnaeve, Ishan Misra, and Nicolas Carion. MDETR – modulated detection for end-to-end multi-modal understanding. In IEEE/CVF International Conference on Computer Vision (ICCV), pages 1780–1790, October 2021. (cited on pp. 3 and 5)
  17. 17.Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, Mustafa Suleyman, and Andrew Zisserman. The kinetics human action video dataset. arXiv preprint arXiv:1705.06950, 2017. (cited on pp. 4 and 6)
  18. 18.Thomas N. Kipf and Max Welling. Semi-supervised classification with graph convolutional networks. In International Conference on Learning Representations, 2017. (cited on p. 2)
  19. 19.Philipp Krähenbühl and Vladlen Koltun. Efficient inference in fully connected CRFs with gaussian edge potentials. In J. Shawe-Taylor, R. Zemel, P. Bartlett, F. Pereira, and K. Q. Weinberger, editors, Advances in Neural Information Processing Systems (NIPS), volume 24. Curran Associates, Inc., 2011. (cited on p. 4)
  20. 20.Harold W. Kuhn. The Hungarian method for the assignment problem. Naval research logistics quarterly, 2(1-2):83–97, 1955. (cited on p. 5)
  21. 21.Chen Liang, Yu Wu, Tianfei Zhou, Wenguan Wang, Zongxin Yang, Yunchao Wei, and Yi Yang. Rethinking cross-modal interaction from a top-down perspective for referring video object segmentation. arXiv preprint arXiv:2106.01061, 2021. (cited on p. 7)
  22. 22.Tsung-Yi Lin, Piotr Dollár, Ross Girshick, Kaiming He, Bharath Hariharan, and Serge Belongie. Feature pyramid networks for object detection. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), July 2017. (cited on p. 4)
  23. 23.Tsung-Yi Lin, Priya Goyal, Ross Girshick, Kaiming He, and Piotr Dollár. Focal loss for dense object detection. In IEEE International Conference on Computer Vision (ICCV), Oct 2017. (cited on p. 5)
  24. 24.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. Microsoft COCO: Common objects in context. In European Conference on Computer Vision (ECCV), pages 740–755. Springer, 2014. (cited on pp. 6 and 7)
  25. 25.Si Liu, Tianrui Hui, Shaofei Huang, Yunchao Wei, Bo Li, and Guanbin Li. Cross-modal progressive comprehension for referring segmentation. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2021. (cited on pp. 2, 4, 6, 7, and 8)
  26. 26.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. RoBERTa: a robustly optimized BERT pretraining approach. arXiv preprint arXiv:1907.11692, 2019. (cited on pp. 2 and 7)
  27. 27.Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. In IEEE/CVF International Conference on Computer Vision (ICCV), pages 10012–10022, October 2021. (cited on pp. 1, 2, 3, and 4)
  28. 28.Ze Liu, Jia Ning, Yue Cao, Yixuan Wei, Zheng Zhang, Stephen Lin, and Han Hu. Video swin transformer. arXiv preprint arXiv:2106.13230, 2021. (cited on pp. 2, 3, 4, and 6)
  29. 29.Junhua Mao, Jonathan Huang, Alexander Toshev, Oana Camburu, Alan L. Yuille, and Kevin Murphy. Generation and comprehension of unambiguous object descriptions. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 2016. (cited on pp. 1 and 7)
  30. 30.Bruce McIntosh, Kevin Duarte, Yogesh S. Rawat, and Mubarak Shah. Visual-textual capsule routing for text-based video segmentation. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2020. (cited on p. 2)
  31. 31.Fausto Milletari, Nassir Navab, and Seyed-Ahmad Ahmadi. V-Net: fully convolutional neural networks for volumetric medical image segmentation. In Fourth international conference on 3D vision (3DV), pages 565–571. IEEE, 2016. (cited on p. 5)
  32. 32.Ke Ning, Lingxi Xie, Fei Wu, and Qi Tian. Polar relative positional encoding for video-language segmentation. In Christian Bessiere, editor, Twenty-Ninth International Joint Conference on Artificial Intelligence, IJCAI-20, pages 948–954. International Joint Conferences on Artificial Intelligence Organization, 7 2020. Main track. (cited on pp. 2 and 4)
  33. 33.Jeffrey Pennington, Richard Socher, and Christopher D. Manning. GloVe: Global vectors for word representation. In Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 1532–1543, Doha, Qatar, Oct. 2014. Association for Computational Linguistics. (cited on p. 7)
  34. 34.Federico Perazzi, Jordi Pont-Tuset, Brian McWilliams, Luc Van Gool, Markus Gross, and Alexander Sorkine-Hornung. A benchmark dataset and evaluation methodology for video object segmentation. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 2016. (cited on p. 6)
  35. 35.Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9, 2019. (cited on pp. 1 and 2)
  36. 36.Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. Faster R-CNN: towards real-time object detection with region proposal networks. In C. Cortes, N. Lawrence, D. Lee, M. Sugiyama, and R. Garnett, editors, Advances in Neural Information Processing Systems, volume 28. Curran Associates, Inc., 2015. (cited on p. 2)
  37. 37.Sara Sabour, Nicholas Frosst, and Geoffrey E. Hinton. Dynamic routing between capsules. In I. Guyon, U. V. Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett, editors, Advances in Neural Information Processing Systems, volume 30. Curran Associates, Inc., 2017. (cited on p. 2)
  38. 38.Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter. arXiv preprint arXiv:1910.01108, 2019. (cited on p. 7)
  39. 39.Seonguk Seo, Joon-Young Lee, and Bohyung Han. URVOS: Unified referring video object segmentation network with a large-scale benchmark. In Andrea Vedaldi, Horst Bischof, Thomas Brox, and Jan-Michael Frahm, editors, European Conference on Computer Vision (ECCV), pages 208–223. Springer, August 2020. (cited on pp. 2, 6, 7, and 8)
  40. 40.Zhi Tian, Chunhua Shen, and Hao Chen. Conditional convolutions for instance segmentation. In Andrea Vedaldi, Horst Bischof, Thomas Brox, and Jan-Michael Frahm, editors, European Conference on Computer Vision (ECCV), pages 213–229. Springer, August 2020. (cited on p. 4)
  41. 41.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In I. Guyon, U. V. Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett, editors, Advances in Neural Information Processing Systems, volume 30. Curran Associates, Inc., 2017. (cited on pp. 1, 2, and 3)
  42. 42.Hao Wang, Cheng Deng, Fan Ma, and Yi Yang. Context modulated dynamic networks for actor and action video segmentation with language queries. AAAI Conference on Artificial Intelligence, 34(07):12152–12159, Apr. 2020. (cited on p. 2)
  43. 43.Hao Wang, Cheng Deng, Junchi Yan, and Dacheng Tao. Asymmetric cross-guided attention network for actor and action video segmentation from natural language query. In IEEE/CVF International Conference on Computer Vision (ICCV), October 2019. (cited on p. 2)
  44. 44.Huiyu Wang, Yukun Zhu, Hartwig Adam, Alan Yuille, and Liang-Chieh Chen. MaX-DeepLab: end-to-end panoptic segmentation with mask transformers. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5463–5474, 2021. (cited on p. 4)
  45. 45.Yuqing Wang, Zhaoliang Xu, Xinlong Wang, Chunhua Shen, Baoshan Cheng, Hao Shen, and Huaxia Xia. End-to-end video instance segmentation with transformers. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 8741–8750, June 2021. (cited on pp. 2, 4, 5, and 7)
  46. 46.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Perric Cistac, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander M. Rush. Transformers: State-of-the-art natural language processing. In Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 38–45, Online, Oct. 2020. Association for Computational Linguistics. (cited on p. 7)
  47. 47.Chenliang Xu, Shao-Hang Hsieh, Caiming Xiong, and Jason J. Corso. Can humans fly? Action understanding with multiple classes of actors. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 2264–2273, 2015. (cited on p. 5)
  48. 48.Ning Xu, Linjie Yang, Yuchen Fan, Jianchao Yang, Dingcheng Yue, Yuchen Liang, Brian Price, Scott Cohen, and Thomas Huang. YouTube-VOS: sequence-to-sequence video object segmentation. In European Conference on Computer Vision (ECCV), pages 585–601, September 2018. (cited on pp. 6 and 7)
  49. 49.Jianhua Yang, Yan Huang, Kai Niu, Zhanyu Ma, and Liang Wang. Actor and action modular network for text-based video segmentation. arXiv preprint arXiv:2011.00786, 2020. (cited on pp. 2 and 6)
  50. 50.Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Russ R. Salakhutdinov, and Quoc V. Le. XLNet: generalized autoregressive pretraining for language understanding. In H. Wallach, H. Larochelle, A. Beygelzimer, F. d'Alché-Buc, E. Fox, and R. Garnett, editors, Advances in Neural Information Processing Systems, volume 32. Curran Associates, Inc., 2019. (cited on pp. 1 and 2)
  51. 51.Linwei Ye, Mrigank Rochan, Zhi Liu, Xiaoqin Zhang, and Yang Wang. Referring segmentation in images and videos with cross-modal self-attention network. IEEE Transactions on Pattern Analysis and Machine Intelligence, (01):1–1, jan 2021. (cited on p. 6)
  52. 52.Licheng Yu, Patrick Poirson, Shan Yang, Alexander C. Berg, and Tamara L. Berg. Modeling context in referring expressions. In European Conference on Computer Vision, pages 69–85. Springer, 2016. (cited on pp. 1 and 7)

Citation

MLA
Botach, A., et al. “End-to-End Referring Video Object Segmentation with Multimodal Transformers”. arXiv, 2021, http://arxiv.org/abs/2111.14821v2.
APA
Botach, A., Zheltonozhskii, E., & Baskin, C. (2021). End-to-End Referring Video Object Segmentation with Multimodal Transformers. arXiv. http://arxiv.org/abs/2111.14821v2
Chicago
Botach, A., E. Zheltonozhskii, and C. Baskin. 2021. “End-to-End Referring Video Object Segmentation with Multimodal Transformers”. arXiv. http://arxiv.org/abs/2111.14821v2.
Harvard
Botach, A., Zheltonozhskii, E. and Baskin, C. (2021) “End-to-End Referring Video Object Segmentation with Multimodal Transformers”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2111.14821v2.
Vancouver
1. Botach A, Zheltonozhskii E, Baskin C (2021) End-to-End Referring Video Object Segmentation with Multimodal Transformers. arXiv

BibTeX

@article{botach2021end,
  title = {End-to-End Referring Video Object Segmentation with Multimodal Transformers},
  author = {Botach, Adam and Zheltonozhskii, Evgenii and Baskin, Chaim},
  year = {2021},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2111.14821v2},
  eprint = {2111.14821}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE