mPLUG: Effective and Efficient Vision-Language Learning by Cross-modal Skip-connections

Chenliang LiHaiyang XuJunfeng TianWei WangMing YanBin BiJiabo YeHe ChenGuohai XuZheng Cao

article2022EMNLP304 citations

Introduces a vision-language foundation model that uses cross-modal skip-connections to eliminate computational bottlenecks on long visual sequences and prevent image features from overwhelming linguistic signals during multi-modal fusion.

Listen

Large artificial intelligence models that connect visual data and human language are increasingly critical for automated visual understanding and text generation. However, current vision-language models face two major technical bottlenecks when processing high-resolution images alongside concise text descriptions. First, long visual sequences often drown out shorter text signals, leading to information loss. Second, calculating detailed relationships across all visual elements is computationally expensive, making training and deployment slow and costly.

The article demonstrates and evaluates mPLUG, an artificial intelligence framework designed to make vision-language learning both more accurate and computationally efficient. The primary objective is to resolve information imbalance and reduce processing overhead by introducing a novel cross-modal skip-connection mechanism.

The authors implemented and evaluated mPLUG using 14 million public image-text pairs across standard training objectives, including cross-modal contrastive learning and language generation. Rather than relying on separate object detectors, the framework extracts image patch representations and text embeddings separately. It then processes them through alternating layers: efficient asymmetric layers that inject visual context into language, followed by unified layers that merge full text and visual streams via skip-connections. The system was benchmarked across standard vision-language benchmarks as well as zero-shot video evaluation tasks without video-specific fine-tuning.

The evaluation produced several significant findings. First, the cross-modal skip-connection network achieved at least a fourfold speedup in cross-modal fusion compared to traditional attention-based networks while improving task performance. Second, on Visual Question Answering benchmarks, mPLUG achieved a score of 81.27, outperforming leading foundation models trained on 60 to 100 times more data. Third, on image captioning benchmarks, it achieved top performance, including a 5.5-point gain in caption accuracy metrics on the MS COCO benchmark over prior state-of-the-art models. Finally, the model demonstrated strong zero-shot transfer capabilities across unseen tasks, outperforming models trained directly on supervised video data on benchmarks like MSR-VTT video retrieval.

These findings indicate that architectural innovation in information fusion can overcome the need for brute-force data scaling. Organizations can achieve superior performance on complex visual and linguistic tasks with significantly less training data and lower computational infrastructure costs, reducing deployment risks and operational budgets.

Decision-makers should consider adopting skip-connected architectures when deploying multimodal intelligence pipelines, particularly in scenarios requiring high image resolutions or low-latency inference. Before full-scale industrial adoption, teams should conduct internal pilot evaluations to determine how the model handles domain-specific vocabularies and obscured visual targets.

Confidence in the reported benchmarks is high given rigorous comparative testing against established baseline models. However, limitations remain: the framework has not yet been scaled to billions of data points or multi-modal mixtures involving unaligned single-modality data, and very long visual sequences combined with extremely short text can still present subtle information-loss challenges.

Cover for mPLUG: Effective and Efficient Vision-Language Learning by Cross-modal Skip-connections

Abstract

Large-scale pre-trained foundation models have been an emerging paradigm for building artificial intelligence (AI) systems, which can be quickly adapted to a wide range of downstream tasks. This paper presents mPLUG, a new vision-language foundation model for both cross-modal understanding and generation. Most existing pre-trained models suffer from inefficiency and linguistic signal overwhelmed by long visual sequences in cross-modal alignment. To address both problems, mPLUG introduces an effective and efficient vision-language architecture with novel cross-modal skip-connections.

mPLUG is pre-trained end-to-end on large-scale image-text pairs with both discriminative and generative objectives. It achieves state-of-the-art results on a wide range of vision-language downstream tasks, including image captioning, image-text retrieval, visual grounding and visual question answering. mPLUG also demonstrates strong zero-shot transferability on vision-language and video-language tasks. The code and pre-trained models are available at https://github.com/alibaba/AliceMind.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Vision-Language Pre-training
  • 2.2 Skip-connection
  • 3 mPLUG
  • 3.1 Model Architecture
  • 3.2 Cross-modal Skip-connected Network
  • 3.3 Pre-training Tasks
  • 4 Experiments
  • 4.1 Data & Setup
  • 4.2 Evaluation on Vision-Language Tasks
  • 4.3 Effectiveness and Efficiency
  • 4.3.1 Analysis of Stride for Skip
  • 4.3.2 Analysis of Cross-modal Fusion
  • 5 Case Study
  • 5.1 Visual Grounding
  • 5.2 VQA
  • 5.3 Zero-shot Transferability
  • 6 Conclusion
  • 7 Limitations
  • References
  • A Implementation Details
  • A.1 Pre-training Tasks
  • A.2 Pre-training Dataset
  • A.3 Pre-training Details
  • A.4 Downstream Task Details
  • A.5 Video Pre-training Data
  • B Visualization of Visual Grounding
  • C Ablation Study and Time-consuming
  • D Differences from BLIP/ALBEF
  • E Comparison Methods

Knowls

  1. Knowl 1 — mPLUG Vision-Language Architecture

    model/method

    mPLUG is a multi-modal foundation model designed for cross-modal understanding and generation that mitigates computational inefficiency and information asymmetry (where long visual sequences overwhelm short linguistic representations).

    The architecture consists of four main components:

    1. Unimodal Visual Encoder: A Vision Transformer (ViT-B/16 or ViT-L/14 initialized from CLIP-ViT) operating directly on flattened 2D image patch embeddings with a prepended [CLS][\text{CLS}] token, yielding a sequence of visual embeddings {vcls,v1,v2,…,vj}\{v_{\text{cls}}, v_1, v_2, \dots, v_j\}.
    2. Unimodal Text Encoder: A 6-layer Transformer initialized from the first 6 layers of BERTbase\text{BERT}_{\text{base}}, which maps tokenized text to embeddings {lcls,l1,l2,…,lk}\{l_{\text{cls}}, l_1, l_2, \dots, l_k\}.
    3. Cross-Modal Skip-Connected Network: A fusion module comprising NN skip-connected fusion blocks. In each block, visual and textual features pass through SS asymmetric co-attention layers (updating text representations using visual queries while omitting vision-side cross-attention for computational efficiency), followed by a single connected-attention layer that processes the concatenated visual representation and updated textual features via joint self-attention.
    4. Transformer Decoder: A 12-layer Transformer decoder that takes the fused cross-modal sequence [vN;lN][v^N; l^N] to perform sequence-to-sequence autoregressive text generation.
  2. Knowl 2 — Cross-Modal Skip-Connected Fusion Mechanism

    equation

    The cross-modal skip-connected network consists of NN blocks. Within each block, cross-modal interactions occur across a stride of SS asymmetric co-attention layers followed by one connected-attention layer.

    For each asymmetric co-attention layer s∈{1,…,S}s \in \{1, \dots, S\}, given text feature ls−1l^{s-1} and visual feature vs−1v^{s-1}:

    lSAs=LN(SA(ls−1)+ls−1)l_{\text{SA}}^s = \text{LN}(\text{SA}(l^{s-1}) + l^{s-1})

    lCAs=LN(CA(lSAs,vs−1)+lSAs)l_{\text{CA}}^s = \text{LN}(\text{CA}(l_{\text{SA}}^s, v^{s-1}) + l_{\text{SA}}^s)

    ls=LN(FFN(lCAs)+lCAs)l^s = \text{LN}(\text{FFN}(l_{\text{CA}}^s) + l_{\text{CA}}^s)

    where SA\text{SA} denotes self-attention, CA(Q,K,V)\text{CA}(Q, K, V) denotes cross-attention with queries QQ derived from text and keys/values K,VK, V derived from visual features vs−1v^{s-1}, FFN\text{FFN} is a feed-forward network, and LN\text{LN} denotes layer normalization. Visual features bypass cross-attention and self-attention in these SS layers.

    At the (S+1)(S+1)-th step of block nn, the original visual feature vn−1v^{n-1} (equivalent to vs−1v^{s-1}) and the co-attended text representation ln−1l^{n-1} (equivalent to lSl^S) are concatenated and passed through a connected-attention layer:

    [vSAn;lSAn]=LN(SA([vn−1;ln−1])+[vn−1;ln−1])[v_{\text{SA}}^n; l_{\text{SA}}^n] = \text{LN}(\text{SA}([v^{n-1}; l^{n-1}]) + [v^{n-1}; l^{n-1}])

    [vn;ln]=LN(FFN([vSAn;lSAn])+[vSAn;lSAn])[v^n; l^n] = \text{LN}(\text{FFN}([v_{\text{SA}}^n; l_{\text{SA}}^n]) + [v_{\text{SA}}^n; l_{\text{SA}}^n])

    The combined output [vn;ln][v^n; l^n] is fed into the subsequent skip-connected block, repeating NN times until producing final representations [vN;lN][v^N; l^N].

  3. Knowl 3 — mPLUG Pre-training Objectives

    model/method

    mPLUG is pre-trained end-to-end on image-text pairs using four joint objectives covering cross-modal understanding and generation:

    1. Image-Text Contrastive Learning (ITC): Applied on unimodal encoder representations before fusion. Computes softmax-normalized bidirectional similarity (image-to-text and text-to-image) using the [CLS][\text{CLS}] embeddings, maintaining two dynamic memory queues of size 65,536 with a momentum coefficient of 0.995.
    2. Image-Text Matching (ITM): A binary classification loss predicting whether an image and text pair match based on the fused multimodal [CLS][\text{CLS}] representation. Hard negative pairs are sampled based on highest contrastive similarity from ITC.
    3. Masked Language Modeling (MLM): Randomly masks 15% of textual tokens and trains the model to predict the masked tokens using the cross-modal skip-connected representations.
    4. Prefix Language Modeling (PrefixLM): Given the connected cross-modal representation of an image and a prefix sub-sequence of the text, the Transformer decoder autoregressively maximizes the likelihood of predicting the remaining text tokens via cross-entropy loss.
  4. Knowl 4 — mPLUG Pre-training Setup and Configuration

    experimental setup

    mPLUG is pre-trained on a corpus of 14M image-text pairs comprising two in-domain datasets (MS COCO: 113K images / 567K texts; Visual Genome: 100K images / 769K texts) and three web datasets (Conceptual Captions 3M: 3M images / 3M texts; Conceptual 12M: 10M images / 10M texts; SBU Captions: 860K images / 860K texts).

    Key configuration parameters include:

    • Architecture Sizes: Text encoder has 6 layers initialized from BERTbase\text{BERT}_{\text{base}} (layers 1-6); skip-connected network has 6 layers initialized from BERTbase\text{BERT}_{\text{base}} (layers 7-12); decoder has 12 Transformer layers. Visual encoder uses CLIP-ViT-B/16 or CLIP-ViT-L/14.
    • Optimization: AdamW with weight decay 0.02, total batch size 1024, 30 epochs on 16 NVIDIA A100 (80GB) GPUs.
    • Learning Rate: Cosine decay with 1000-iteration warmup. For mPLUG (ViT-B), peak rates are 1e-5 (ViT) and 1e-4 (BERTbase\text{BERT}_{\text{base}}). For mPLUG (ViT-L), peak rates are 5e-6 (ViT) and 5e-5 (BERTbase\text{BERT}_{\text{base}}), decaying to 1e-6.
    • Input Resolutions: Pre-training uses random crops of 256×256256 \times 256 (ViT-B) or 224×224224 \times 224 (ViT-L) with RandAugment. Downstream fine-tuning increases resolution to 336×336336 \times 336 (or 504×504504 \times 504 for VQA).
  5. Knowl 5 — Downstream Performance on VQA, Image Captioning, and NoCaps

    data/table

    mPLUG was evaluated on Visual Question Answering (VQA v2.0, evaluated via unconstrained open-vocabulary generation), MS COCO Caption (Karpathy test split, under cross-entropy and CIDEr optimization), and NoCaps validation set.

    Models # Data VQA COCO (Cross-Entropy) COCO (CIDEr Opt.) NoCaps
    Std Dev B@4 M C S B@4 M C S C
    VinVL 5.65M 76.52 76.60 38.5 30.4 130.8 23.4 41.0 31.1 140.9 25.2 97.3
    BLIP 129M 78.25 78.32 40.4 - 136.7 - - - - - 113.2
    VLMo - 79.94 79.98 - - - - - - - - -
    OFA 18M 79.87 80.02 - - - - 43.5 31.9 149.6 26.1 -
    SimVLMlarge_{\text{large}} 1.8B 80.03 80.34 40.3 33.4 142.6 24.7 - - - - -
    Florence 0.9B 80.16 80.36 - - - - - - - - -
    GIT 0.8B 78.81 - 44.1 31.5 144.8 24.7 44.1 32.2 151.1 26.3 125.5
    mPLUGViT-B_{\text{ViT-B}} 14M 79.89 79.92 41.5 31.1 137.5 23.8 44.9 31.2 150.4 25.2 108.5
    mPLUGViT-L_{\text{ViT-L}} 14M 81.27 81.26 43.1 31.4 141.0 24.2 46.5 32.0 155.1 26.0 114.8

    With 14M pre-training pairs, mPLUG (ViT-L) achieves an accuracy of 81.27% on VQA Test-std, outperforming SimVLM (1.8B pairs) and Florence (0.9B pairs), and achieves a CIDEr score of 155.1 on COCO Caption under CIDEr optimization.

  6. Knowl 6 — Image-Text Retrieval Results on MSCOCO and Flickr30K

    data/table

    Cross-modal retrieval performance for Text Retrieval (TR) and Image Retrieval (IR) evaluated on MSCOCO (5K test split) and Flickr30K (1K test split):

    Models Data MSCOCO TR MSCOCO IR Flickr30K TR Flickr30K IR
    R@1 R@5 R@10 R@1 R@5 R@10 R@1 R@5 R@10 R@1 R@5 R@10
    UNITER 4M 65.7 88.6 93.8 52.9 79.9 88.0 87.3 98.0 99.2 75.6 94.1 96.8
    OSCAR 4M 70.0 91.1 95.5 54.0 80.8 88.5 - - - - - -
    VLMo 4M 78.2 94.4 97.4 60.6 84.4 91.0 95.3 99.9 100.0 84.5 97.3 98.6
    ALIGN 1.8B 77.0 93.5 96.9 59.9 83.3 89.8 95.3 99.8 100.0 84.9 97.4 98.6
    ALBEF 14M 77.6 94.3 97.2 60.7 84.3 90.5 95.9 99.8 100.0 85.6 97.5 98.9
    Florence 0.9B 81.8 95.2 - 63.2 85.7 - 97.2 99.9 - 87.9 98.1 -
    BLIP 14M 80.6 95.2 97.6 63.1 85.3 91.1 96.6 99.8 100.0 87.2 97.5 98.8
    BLIP 129M 82.4 95.4 97.9 65.1 86.3 91.8 97.4 99.8 99.9 87.6 97.7 99.0
    mPLUG 14M 82.8 96.1 98.3 65.8 87.3 92.6 97.6 100.0 100.0 88.4 97.9 99.1

    mPLUG with 14M pre-training images achieves superior recall across all metrics compared to models trained on orders of magnitude more data (e.g., BLIP with 129M images and Florence with 0.9B images), outperforming 14M-trained BLIP by +2.2% TR R@1 on MSCOCO and +1.0% TR R@1 on Flickr30K.

  7. Knowl 7 — Visual Grounding, NLVR2, and SNLI-VE Performance

    data/table

    Evaluation of mPLUG on referring expression grounding (accuracy with IoU≥0.5\text{IoU} \ge 0.5 on RefCOCO, RefCOCO+, and RefCOCOg), visual reasoning on NLVR2, and visual entailment on SNLI-VE:

    Model RefCOCO RefCOCO+ RefCOCOg NLVR2 SNLI-VE
    val testA testB val testA testB val-u test-u test-P test
    UNITER 81.41 87.04 74.17 75.90 81.45 66.70 74.86 75.77 79.98 79.38
    METER - - - - - - - - 83.05 81.19
    ALBEF - - - - - - - - 83.14 80.91
    VILLA 82.39 87.48 74.84 76.17 81.54 66.84 76.18 76.71 81.47 80.02
    MDETR 86.75 89.58 81.41 79.52 84.09 70.62 81.64 80.89 - -
    UNICORN 88.29 90.42 83.06 80.30 85.05 71.88 83.44 83.93 - -
    VLMo - - - - - - - - 86.86 -
    SimVLMlarge_{\text{large}} - - - - - - - - 84.84 85.62
    OFA 90.05 92.93 85.26 84.49 90.10 77.77 84.54 85.20 - 90.20
    mPLUG 92.40 94.51 88.42 86.02 90.17 78.17 85.88 86.42 84.95 89.29

    For visual grounding, mPLUG feeds concatenated visual and attended textual representations to its decoder to generate bounding box coordinates directly, achieving improvements over OFA on complex splits (+3.16% on RefCOCO testB and +1.22% on RefCOCOg test-u).

  8. Knowl 8 — Effect of Fusion Stride $S$ on Efficiency and Downstream Performance

    empirical result

    The stride parameter SS controls the frequency of full connected-attention layers relative to asymmetric co-attention layers within the cross-modal skip-connected network. Testing mPLUG (ViT-B) across stride values S∈{1,2,3,6}S \in \{1, 2, 3, 6\} on a 6-layer fusion network shows:

    • Computational Speedup: As SS increases from 1 to 6, the forward running time of the skip-connected network for 100 samples decreases from 1.65s to 0.28s, representing an approximate 5.9×5.9\times speedup for the fusion module. On a single NVIDIA V100 GPU, end-to-end forward inference time per 100 samples drops from 5.06s to 3.69s for mPLUG (ViT-B) and from 3.08s to 1.71s for mPLUG (ViT-S).
    • Downstream Accuracy: Downstream performance on VQA test-dev and NLVR2 test-P increases from S=1S=1 to S=3S=3 and stabilizes with a marginal decline from S=3S=3 to S=6S=6. S=6S=6 achieves accuracy comparable to S=3S=3 while providing an overall 30% inference speedup, making S=6S=6 the default setting for mPLUG (ViT-L).
  9. Knowl 9 — Comparison of Cross-Modal Fusion Paradigms and Resolution Scaling

    empirical result

    A comparative analysis of fusion architectures (Connected-Attention, Co-Attention, Asymmetric Co-Attention, and Cross-Modal Skip-Connected) demonstrates two key properties:

    1. Fusion Efficiency & Quality Tradeoff: Connected-attention and bidirectional co-attention require high computation time (forward times of 1.65s and 1.86s for 100 samples). Asymmetric co-attention alone is fast (0.28s) but suffers from lower accuracy on VQA test-dev and NLVR2 test-P due to language bias and forgetting visual representations. The cross-modal skip-connected network achieves the highest accuracy with a 4×4\times overall speedup over connected and bidirectional co-attention.
    2. Robustness to Long Visual Sequences: Lengthening visual sequences via increasing image resolution (from 224 up to 576, where visual sequence length reaches 60×60\times textual sequence length) widens the performance advantage of skip-connected fusion over connected-attention on VQA test-dev. This confirms that skip-connections prevent short linguistic signals from being overwhelmed by long visual sequences during cross-modal alignment.
  10. Knowl 10 — Zero-Shot Transferability to Video-Language Tasks

    data/table

    Without pre-training on video data, mPLUG demonstrates zero-shot transfer to video-language downstream tasks by uniformly sampling frames (n=8n=8 for retrieval and captioning, n=16n=16 for question answering) and concatenating frame features into a single sequence.

    Model Pretrain Data MSRVTT Retrieval (Zero-Shot) MSRVTT-QA VATEX-Cap
    R@1 R@5 R@10 Acc CIDEr
    MIL-NCE How100M 9.9 24.0 32.4 - -
    VideoCLIP How100M 10.4 22.2 30.0 - -
    CLIP WIT400M 26.0 49.4 60.7 - -
    Florence FLD900M 37.6 63.8 72.6 - -
    BLIP 129M 43.3 65.6 74.7 19.2 37.4
    mPLUG (Zero-shot) 14M 38.1 59.2 68.2 - -
    mPLUG†^{\dagger} (COCO FT) 14M 44.3 66.4 75.4 21.1 42.0

    When fine-tuned only on COCO image retrieval, zero-shot mPLUG achieves 44.3% R@1 on MSRVTT 1k test split, outperforming models specifically pre-trained or fine-tuned on dedicated video datasets (e.g., VideoCLIP, Florence, and BLIP).

  11. Knowl 11 — Ablation of Pre-training Tasks in mPLUG

    data/table

    Ablation experiments isolating the impact of individual pre-training objectives on mPLUG (ViT-B) evaluated on the VQA test-dev benchmark:

    Configuration VQA test-dev (%)
    mPLUG_ViT-B 79.89
    w/o ITC 78.17
    w/o PrefixLM 78.45
    w/o ITM 79.36
    w/o MLM 79.54

    Removing Image-Text Contrastive Learning (ITC) causes the largest performance drop (-1.72%), followed by Prefix Language Modeling (PrefixLM, -1.44%), demonstrating that contrastive alignment and generative prefix modeling are the primary contributors to cross-modal representation learning.

  12. Knowl 12 — Stated Limitations of mPLUG

    limitation

    The authors identify three primary limitations of the mPLUG model:

    1. Scalability: Pre-training is restricted to 14M image-text pairs on a 12+12 layer Transformer encoder-decoder; performance scaling behavior on larger corpora with additional single-modality (image-only, text-only) or labeled data remains unverified.
    2. Vision Encoder Dependency: The visual representations rely on a public off-the-shelf CLIP-ViT backbone rather than a custom-trained encoder optimized specifically for semantic feature extraction on large-scale multimodal data.
    3. Residual Information Vanishing: Because connected-attention layers are retained at the end of each skip block, information vanishing caused by long visual sequences still persists when the text sequence is extremely short relative to the visual input.

Coverage note — None. All primary contributions, architecture specifications, mathematical formulations, experimental benchmark results, ablation studies, and stated limitations have been fully captured.

References

  1. 1.Harsh Agrawal, Karan Desai, Yufei Wang, Xinlei Chen, Rishabh Jain, Mark Johnson, Dhruv Batra, Devi Parikh, Stefan Lee, and Peter Anderson. 2018. nocaps: novel object captioning at scale. CoRR, abs/1812.08658.
  2. 2.Hassan Akbari, Liangzhe Yuan, Rui Qian, Wei-Hong Chuang, Shih-Fu Chang, Yin Cui, and Boqing Gong. 2021. Vatt: Transformers for multimodal self-supervised learning from raw video, audio and text. Advances in Neural Information Processing Systems, 34.
  3. 3.Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang. 2018. Bottom-up and top-down attention for image captioning and visual question answering. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 6077–6086.
  4. 4.Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh. 2015. Vqa: Visual question answering. In Proceedings of the IEEE international conference on computer vision, pages 2425–2433.
  5. 5.Max Bain, Arsha Nagrani, Gül Varol, and Andrew Zisserman. 2021. Frozen in time: A joint video and image encoder for end-to-end retrieval. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 1728–1738.
  6. 6.Bin Bi, Chenliang Li, Chen Wu, Ming Yan, Wei Wang, Songfang Huang, Fei Huang, and Luo Si. 2020. Palm: Pre-training an autoencoding&autoregressive language model for context-conditioned generation. arXiv preprint arXiv:2004.07159.
  7. 7.Soravit Changpinyo, Piyush Sharma, Nan Ding, and Radu Soricut. 2021. Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3558–3568.
  8. 8.Xinlei Chen, Hao Fang, Tsung-Yi Lin, Ramakrishna Vedantam, Saurabh Gupta, Piotr Dollár, and C Lawrence Zitnick. 2015a. Microsoft coco captions: Data collection and evaluation server. arXiv preprint arXiv:1504.00325.
  9. 9.Xinlei Chen, Hao Fang, Tsung-Yi Lin, Ramakrishna Vedantam, Saurabh Gupta, Piotr Dollár, and C. Lawrence Zitnick. 2015b. Microsoft COCO captions: Data collection and evaluation server. CoRR, abs/1504.00325.
  10. 10.Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu. 2020. Uniter: Universal image-text representation learning. In European conference on computer vision, pages 104–120. Springer.
  11. 11.Jaemin Cho, Jie Lei, Hao Tan, and Mohit Bansal. 2021. Unifying vision-and-language tasks via text generation. In Proceedings of the 38th International Conference on Machine Learning, volume 139 of Proceedings of Machine Learning Research, pages 1931–1942. PMLR.
  12. 12.Ekin D Cubuk, Barret Zoph, Jonathon Shlens, and Quoc V Le. 2020. Randaugment: Practical automated data augmentation with a reduced search space. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, pages 702–703.
  13. 13.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805.
  14. 14.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. 2020. An image is worth 16x16 words: Transformers for image recognition at scale. In International Conference on Learning Representations.
  15. 15.Zi-Yi Dou, Yichong Xu, Zhe Gan, Jianfeng Wang, Shuohang Wang, Lijuan Wang, Chenguang Zhu, Zicheng Liu, Michael Zeng, et al. 2021. An empirical study of training end-to-end vision-and-language transformers. arXiv preprint arXiv:2111.02387.
  16. 16.Tsu-Jui Fu, Linjie Li, Zhe Gan, Kevin Lin, William Yang Wang, Lijuan Wang, and Zicheng Liu. 2021. Violet: End-to-end video-language transformers with masked visual-token modeling. arXiv preprint arXiv:2111.12681.
  17. 17.Zhe Gan, Yen-Chun Chen, Linjie Li, Chen Zhu, Yu Cheng, and Jingjing Liu. 2020. Large-scale adversarial training for vision-and-language representation learning. In NeurIPS.
  18. 18.Jort F Gemmeke, Daniel PW Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R Channing Moore, Manoj Plakal, and Marvin Ritter. 2017. Audio set: An ontology and human-labeled dataset for audio events. In 2017 IEEE international conference on acoustics, speech and signal processing (ICASSP), pages 776–780. IEEE.
  19. 19.Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh. 2017. Making the v in vqa matter: Elevating the role of image understanding in visual question answering. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 6904–6913.
  20. 20.Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick. 2020. Momentum contrast for unsupervised visual representation learning. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 9729–9738.
  21. 21.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778.
  22. 22.Xiaowei Hu, Zhe Gan, Jianfeng Wang, Zhengyuan Yang, Zicheng Liu, Yumao Lu, and Lijuan Wang. 2021. Scaling up vision-language pre-training for image captioning. CoRR, abs/2111.12233.
  23. 23.Gao Huang, Zhuang Liu, Laurens Van Der Maaten, and Kilian Q Weinberger. 2017. Densely connected convolutional networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 4700–4708.
  24. 24.Zhicheng Huang, Zhaoyang Zeng, Bei Liu, Dongmei Fu, and Jianlong Fu. 2020. Pixel-bert: Aligning image pixels with text by deep multi-modal transformers. arXiv preprint arXiv:2004.00849.
  25. 25.Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc V Le, Yunhsuan Sung, Zhen Li, and Tom Duerig. 2021. Scaling up visual and vision-language representation learning with noisy text supervision. arXiv preprint arXiv:2102.05918.
  26. 26.Aishwarya Kamath, Mannat Singh, Yann LeCun, Gabriel Synnaeve, Ishan Misra, and Nicolas Carion. 2021. Mdetr-modulated detection for end-to-end multi-modal understanding. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 1780–1790.
  27. 27.Andrej Karpathy and Li Fei-Fei. 2015. Deep visual-semantic alignments for generating image descriptions. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 3128–3137.
  28. 28.Wonjae Kim, Bokyung Son, and Ildoo Kim. 2021. Vilt: Vision-and-language transformer without convolution or region supervision. arXiv preprint arXiv:2102.03334.
  29. 29.Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A Shamma, et al. 2017. Visual genome: Connecting language and vision using crowdsourced dense image annotations. International journal of computer vision, 123(1):32–73.
  30. 30.Dongxu Li, Junnan Li, Hongdong Li, Juan Carlos Niebles, and Steven CH Hoi. 2021a. Align and prompt: Video-and-language pre-training with entity prompts. arXiv preprint arXiv:2112.09583.
  31. 31.Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi. 2022. Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation. arXiv preprint arXiv:2201.12086.
  32. 32.Junnan Li, Ramprasaath Selvaraju, Akhilesh Gotmare, Shafiq Joty, Caiming Xiong, and Steven Chu Hong Hoi. 2021b. Align before fuse: Vision and language representation learning with momentum distillation. Advances in Neural Information Processing Systems, 34.
  33. 33.Liunian Harold Li, Mark Yatskar, Da Yin, Cho-Jui Hsieh, and Kai-Wei Chang. 2019. Visualbert: A simple and performant baseline for vision and language. arXiv preprint arXiv:1908.03557.
  34. 34.Wei Li, Can Gao, Guocheng Niu, Xinyan Xiao, Hao Liu, Jiachen Liu, Hua Wu, and Haifeng Wang. 2020a. Unimo: Towards unified-modal understanding and generation via cross-modal contrastive learning. arXiv preprint arXiv:2012.15409.
  35. 35.Xiujun Li, Xi Yin, Chunyuan Li, Pengchuan Zhang, Xiaowei Hu, Lei Zhang, Lijuan Wang, Houdong Hu, Li Dong, Furu Wei, et al. 2020b. Oscar: Object-semantics aligned pre-training for vision-language tasks. In European Conference on Computer Vision, pages 121–137. Springer.
  36. 36.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. 2014. Microsoft coco: Common objects in context. In European conference on computer vision, pages 740–755. Springer.
  37. 37.Fenglin Liu, Xuancheng Ren, Zhiyuan Zhang, Xu Sun, and Yuexian Zou. 2021. Rethinking skip connection with layer normalization in transformers and resnets. arXiv preprint arXiv:2105.07205.
  38. 38.Ilya Loshchilov and Frank Hutter. 2017. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101.
  39. 39.Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. 2019. Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks. In Advances in Neural Information Processing Systems, pages 13–23.
  40. 40.Junhua Mao, Jonathan Huang, Alexander Toshev, Oana Camburu, Alan L Yuille, and Kevin Murphy. 2016. Generation and comprehension of unambiguous object descriptions. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 11–20.
  41. 41.Antoine Miech, Jean-Baptiste Alayrac, Lucas Smaira, Ivan Laptev, Josef Sivic, and Andrew Zisserman. 2020. End-to-end learning of visual representations from uncurated instructional videos. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9879–9889.
  42. 42.Antoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi, Ivan Laptev, and Josef Sivic. 2019. Howto100m: Learning a text-video embedding by watching hundred million narrated video clips. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 2630–2640.
  43. 43.Vicente Ordonez, Girish Kulkarni, and Tamara L Berg. 2011. Im2text: Describing images using 1 million captioned photographs. In Advances in neural information processing systems, pages 1143–1151.
  44. 44.Bryan A Plummer, Liwei Wang, Chris M Cervantes, Juan C Caicedo, Julia Hockenmaier, and Svetlana Lazebnik. 2015. Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models. In Proceedings of the IEEE international conference on computer vision, pages 2641–2649.
  45. 45.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. arXiv preprint arXiv:2103.00020.
  46. 46.Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. 2022. Hierarchical text-conditional image generation with clip latents. arXiv preprint arXiv:2204.06125.
  47. 47.Steven J. Rennie, Etienne Marcheret, Youssef Mroueh, Jerret Ross, and Vaibhava Goel. 2017. Self-critical sequence training for image captioning. In 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 1179–1195.
  48. 48.Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut. 2018. Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2556–2565.
  49. 49.Sheng Shen, Liunian Harold Li, Hao Tan, Mohit Bansal, Anna Rohrbach, Kai-Wei Chang, Zhewei Yao, and Kurt Keutzer. 2021. How much can clip benefit vision-and-language tasks? arXiv preprint arXiv:2107.06383.
  50. 50.Rupesh Kumar Srivastava, Klaus Greff, and Jürgen Schmidhuber. 2015. Highway networks. arXiv preprint arXiv:1505.00387.
  51. 51.Weijie Su, Xizhou Zhu, Yue Cao, Bin Li, Lewei Lu, Furu Wei, and Jifeng Dai. 2019. Vl-bert: Pre-training of generic visual-linguistic representations. arXiv preprint arXiv:1908.08530.
  52. 52.Alane Suhr, Stephanie Zhou, Ally Zhang, Iris Zhang, Huajun Bai, and Yoav Artzi. 2018. A corpus for reasoning about natural language grounded in photographs. arXiv preprint arXiv:1811.00491.
  53. 53.Hao Tan and Mohit Bansal. 2019. Lxmert: Learning cross-modality encoder representations from transformers. arXiv preprint arXiv:1908.07490.
  54. 54.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in neural information processing systems, pages 5998–6008.
  55. 55.Jianfeng Wang, Xiaowei Hu, Zhe Gan, Zhengyuan Yang, Xiyang Dai, Zicheng Liu, Yumao Lu, and Lijuan Wang. 2021a. UFO: A unified transformer for vision-language representation learning. CoRR, abs/2111.10023.
  56. 56.Jianfeng Wang, Zhengyuan Yang, Xiaowei Hu, Linjie Li, Kevin Lin, Zhe Gan, Zicheng Liu, Ce Liu, and Lijuan Wang. 2022a. Git: A generative image-to-text transformer for vision and language. arXiv preprint arXiv:2205.14100.
  57. 57.Peng Wang, An Yang, Rui Men, Junyang Lin, Shuai Bai, Zhikang Li, Jianxin Ma, Chang Zhou, Jingren Zhou, and Hongxia Yang. 2022b. Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework. arXiv preprint arXiv:2202.03052.
  58. 58.Wei Wang, Bin Bi, Ming Yan, Chen Wu, Zuyi Bao, Jiangnan Xia, Liwei Peng, and Luo Si. 2019. Structbert: incorporating language structures into pre-training for deep language understanding. arXiv preprint arXiv:1908.04577.
  59. 59.Wenhui Wang, Hangbo Bao, Li Dong, and Furu Wei. 2021b. Vlmo: Unified vision-language pre-training with mixture-of-modality-experts. arXiv preprint arXiv:2111.02358.
  60. 60.Xinyu Wang, Jiong Cai, Yong Jiang, Pengjun Xie, Kewei Tu, and Wei Lu. 2022c. Named entity and relation extraction with multi-modal retrieval. In Proceedings of EMNLP.
  61. 61.Xinyu Wang, Min Gui, Yong Jiang, Zixia Jia, Nguyen Bach, Tao Wang, Zhongqiang Huang, and Kewei Tu. 2022d. ITA: Image-text alignments for multi-modal named entity recognition. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 3176–3189, Seattle, United States. Association for Computational Linguistics.
  62. 62.Zirui Wang, Jiahui Yu, Adams Wei Yu, Zihang Dai, Yulia Tsvetkov, and Yuan Cao. 2021c. Simvlm: Simple visual language model pretraining with weak supervision. CoRR, abs/2108.10904.
  63. 63.Ning Xie, Farley Lai, Derek Doran, and Asim Kadav. 2019. Visual entailment: A novel task for fine-grained image understanding. CoRR, abs/1901.06706.
  64. 64.Haiyang Xu, Ming Yan, Chenliang Li, Bin Bi, Songfang Huang, Wenming Xiao, and Fei Huang. 2021a. E2e-vlp: End-to-end vision-language pre-training enhanced by visual learning. arXiv preprint arXiv:2106.01804.
  65. 65.Hu Xu, Gargi Ghosh, Po-Yao Huang, Dmytro Okhonko, Armen Aghajanyan, Florian Metze, Luke Zettlemoyer, and Christoph Feichtenhofer. 2021b. Videoclip: Contrastive pre-training for zero-shot video-text understanding. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 6787–6800.
  66. 66.Ming Yan, Haiyang Xu, Chenliang Li, Junfeng Tian, Bin Bi, Wei Wang, Weihua Chen, Xianzhe Xu, Fan Wang, Zheng Cao, Zhicheng Zhang, Qiyu Zhang, Ji Zhang, Songfang Huang, Fei Huang, Luo Si, and Rong Jin. 2021. Achieving human parity on visual question answering. CoRR, abs/2111.08896.
  67. 67.Antoine Yang, Antoine Miech, Josef Sivic, Ivan Laptev, and Cordelia Schmid. 2021a. Just ask: Learning to answer questions from millions of narrated videos. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 1686–1697.
  68. 68.Zhengyuan Yang, Zhe Gan, Jianfeng Wang, Xiaowei Hu, Faisal Ahmed, Zicheng Liu, Yumao Lu, and Lijuan Wang. 2021b. Crossing the format boundary of text and boxes: Towards unified vision-language modeling. CoRR, abs/2111.12085.
  69. 69.Lewei Yao, Runhui Huang, Lu Hou, Guansong Lu, Minzhe Niu, Hang Xu, Xiaodan Liang, Zhenguo Li, Xin Jiang, and Chunjing Xu. 2021. Filip: Fine-grained interactive language-image pre-training. arXiv preprint arXiv:2111.07783.
  70. 70.Fei Yu, Jiji Tang, Weichong Yin, Yu Sun, Hao Tian, Hua Wu, and Haifeng Wang. 2021. Ernie-vil: Knowledge enhanced vision-language representations through scene graphs. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, pages 3208–3216.
  71. 71.Licheng Yu, Patrick Poirson, Shan Yang, Alexander C Berg, and Tamara L Berg. 2016. Modeling context in referring expressions. In European Conference on Computer Vision, pages 69–85. Springer.
  72. 72.Lu Yuan, Dongdong Chen, Yi-Ling Chen, Noel Codella, Xiyang Dai, Jianfeng Gao, Houdong Hu, Xuedong Huang, Boxin Li, Chunyuan Li, et al. 2021. Florence: A new foundation model for computer vision. arXiv preprint arXiv:2111.11432.
  73. 73.Rowan Zellers, Ximing Lu, Jack Hessel, Youngjae Yu, Jae Sung Park, Jize Cao, Ali Farhadi, and Yejin Choi. 2021. Merlot: Multimodal neural script knowledge models. Advances in Neural Information Processing Systems, 34.
  74. 74.Pengchuan Zhang, Xiujun Li, Xiaowei Hu, Jianwei Yang, Lei Zhang, Lijuan Wang, Yejin Choi, and Jianfeng Gao. 2021. Vinvl: Making visual representations matter in vision-language models. CoRR, abs/2101.00529.

Citation

MLA
Li, C., et al. “mPLUG: Effective and Efficient Vision-Language Learning by Cross-modal Skip-connections”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 2022, pp. 7241–59, https://doi.org/10.18653/v1/2022.emnlp-main.488.
APA
Li, C., Xu, H., Tian, J., Wang, W., Yan, M., Bi, B., Ye, J., Chen, H., Xu, G., Cao, Z., Zhang, J., Huang, S., Huang, F., Zhou, J., & Si, L. (2022). mPLUG: Effective and Efficient Vision-Language Learning by Cross-modal Skip-connections. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 7241–7259. https://doi.org/10.18653/v1/2022.emnlp-main.488
Chicago
Li, C., H. Xu, J. Tian, et al. 2022. “mPLUG: Effective and Efficient Vision-Language Learning by Cross-modal Skip-connections”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 7241–59. https://doi.org/10.18653/v1/2022.emnlp-main.488.
Harvard
Li, C. et al. (2022) “mPLUG: Effective and Efficient Vision-Language Learning by Cross-modal Skip-connections”, Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 7241–7259. Available at: https://doi.org/10.18653/v1/2022.emnlp-main.488.
Vancouver
1. Li C, Xu H, Tian J, et al (2022) mPLUG: Effective and Efficient Vision-Language Learning by Cross-modal Skip-connections. In: Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 7241–7259

BibTeX

@inproceedings{li-etal-2022-mplug,
    title = "m{PLUG}: Effective and Efficient Vision-Language Learning by Cross-modal Skip-connections",
    author = "Li, Chenliang  and
      Xu, Haiyang  and
      Tian, Junfeng  and
      Wang, Wei  and
      Yan, Ming  and
      Bi, Bin  and
      Ye, Jiabo  and
      Chen, He  and
      Xu, Guohai  and
      Cao, Zheng  and
      Zhang, Ji  and
      Huang, Songfang  and
      Huang, Fei  and
      Zhou, Jingren  and
      Si, Luo",
    editor = "Goldberg, Yoav  and
      Kozareva, Zornitsa  and
      Zhang, Yue",
    booktitle = "Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2022",
    address = "Abu Dhabi, United Arab Emirates",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.emnlp-main.488/",
    doi = "10.18653/v1/2022.emnlp-main.488",
    pages = "7241--7259"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/