MultiInstruct: Improving Multi-Modal Zero-Shot Learning via Instruction Tuning

Zhiyang XuYing ShenLifu Huang

article2023ACL145 citationsOutstanding Paper Award

Introduces the first multimodal instruction tuning benchmark spanning 62 diverse tasks across 10 categories to boost zero-shot generalization and reduce sensitivity to prompt phrasing in vision-language models.

Listen

Instruction tuning has enabled large language models to generalize to unseen natural language tasks without task-specific training, but its potential for multimodal systems remains largely unexplored. The article addresses this gap by introducing MULTIINSTRUCT, the first benchmark dataset designed for multimodal instruction tuning. The primary objective is to demonstrate that fine-tuning vision-language models on diverse multimodal tasks guided by natural language instructions significantly improves their zero-shot generalization to unseen tasks, while reducing sensitivity to variations in prompt wording.

To evaluate this approach, the authors compiled 62 multimodal tasks spanning 10 broad categories derived from 21 open-source datasets, pairing each task with five expert-written instruction templates. All tasks were mapped into a unified sequence-to-sequence format where text, images, and visual coordinates share a single vocabulary. Using the pre-trained OFA model as a base, the study fine-tuned the system on 53 tasks and evaluated zero-shot performance across nine unseen multimodal tasks and 20 text-only tasks. The authors also explored transfer learning strategies using NATURAL INSTRUCTIONS, a large-scale repository of 832 English text-only tasks, and introduced a Sensitivity metric to measure model output stability across different prompt wordings.

The findings show that multimodal instruction tuning substantially improves zero-shot capabilities. For instance, on Grounded Visual Question Answering, the baseline model scored near zero because it failed to follow spatial instructions, whereas the instruction-tuned model achieved an average accuracy of 47.22%. Training on multiple diverse instructions reduced model sensitivity to prompt variations by more than half compared to the base model. Notably, tuning exclusively on text-only instructions degraded vision-language performance because the attention layers shifted focus away from visual tokens. However, combining text-only and multimodal tasks simultaneously achieved the strongest overall performance, preserving text reasoning while improving multimodal stability.

These results demonstrate that multimodal instruction tuning is a practical and robust strategy for building flexible, generalist artificial intelligence systems that do not require custom re-training for new use cases. For organizations deploying vision-language models, the authors recommend incorporating instruction diversity during training and adopting mixed multimodal-text datasets rather than sequential or single-modality tuning. Future work should address current limitations by expanding beyond English datasets, incorporating audio and video modalities, and testing larger foundational models.

arXiv: 2212.10773VT-NLP/MultiInstruct
Cover for MultiInstruct: Improving Multi-Modal Zero-Shot Learning via Instruction Tuning

Abstract

Instruction tuning, a new learning paradigm that fine-tunes pre-trained language models on tasks specified through instructions, has shown promising zero-shot performance on various natural language processing tasks. However, it has yet to be explored for vision and multimodal tasks. In this work, we introduce MULTIINSTRUCT, the first multimodal instruction tuning benchmark dataset that consists of 62 diverse multimodal tasks in a unified seq-to-seq format covering 10 broad categories. The tasks are derived from 21 existing open-source datasets and each task is equipped with 5 expert-written instructions. We take OFA (Wang et al., 2022a) as the base pre-trained model for multimodal instruction tuning, and to further improve its zero-shot performance, we explore multiple transfer learning strategies to leverage the large-scale NATURAL INSTRUCTIONS dataset (Mishra et al., 2022). Experimental results demonstrate strong zero-shot performance on various unseen multimodal tasks and the benefit of transfer learning from a text-only instruction dataset. We also design a new evaluation metric – Sensitivity, to evaluate how sensitive the model is to the variety of instructions. Our results indicate that fine-tuning the model on a diverse set of tasks and instructions leads to a reduced sensitivity to variations in instructions for each task1.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 MULTIINSTRUCT
  • 3.1 Multimodal Task and Data Collection
  • 3.2 Task Instruction Creation
  • 3.3 Multimodal Instruction Formatting
  • 4 Problem Setup and Models
  • 4.1 Problem Setup
  • 4.2 Transfer Learning from NATURAL INSTRUCTIONS
  • 5 Experimental Setup
  • 6 Results and Discussion
  • 6.1 Effectiveness of Instruction Tuning on MULTIINSTRUCT
  • 6.2 Impact of Transfer Learning from NATURAL INSTRUCTIONS
  • 6.3 Impact of Increasing Multimodal Instruction Task Clusters
  • 6.4 Effect of Diverse Instructions on Instruction Tuning
  • 6.5 Effect of Fine-tuning Strategies on Model Sensitivity
  • 7 Zero-Shot Performance on NLP Tasks
  • 8 Conclusion
  • Limitations
  • Limitations of Experiments and Evaluation
  • Acknowledgments
  • References
  • A Tasks Defined in MULTIINSTRUCT
  • B More Details for Experimental Setup
  • B.1 Multimodal Evaluation Datasets
  • B.2 NLP Evaluation Tasks
  • B.3 Approaches for Comparison
  • B.4 Training Details
  • C Attention Analysis

Knowls

  1. Knowl 1 — MultiInstruct Multimodal Instruction Tuning Benchmark

    definition

    MultiInstruct is a benchmark dataset designed for multimodal instruction tuning across vision-language and multimodal reasoning tasks. It consists of 62 distinct multimodal tasks organized into 10 broad categories:

    1. Visual Question Answering (VQA) (e.g., Open-Domain VQA, Compositional VQA, Outside Knowledge VQA, Text VQA, Grounded VQA)
    2. Commonsense Reasoning (e.g., Visual Spatial Reasoning, Visual Entailment, Natural Language for Visual Reasoning, Commonsense VQA)
    3. Grounded Generation (e.g., Grounded Captioning, Visual Grounding, Object Grounding, Text Localization, Referring Expression Grounding/Generation)
    4. Grounded Matching (e.g., Region-Caption Matching, Grounded Caption Selection, Object-Region Selection, Referring Expression Selection)
    5. Region Understanding (e.g., Most/Least-overlapping Region Selection, Overlapping Region Selection, Region Area, Region Overlapping Detection)
    6. Image Understanding (e.g., Color/Scene/Object Recognition, Counting, Sentiment Understanding, Position Reasoning, Utility Affordance, Image Quality)
    7. Visual Relationship (e.g., Object Relationship, Visual Subject/Object Identification, Visual Subject/Object Localization, Grounded Image Attribute Identification)
    8. Image-Text Matching (e.g., Image-Text Matching, Question-Image Matching, Image-Text Selection)
    9. Temporal Ordering (e.g., WikiHow Next Step Generation/Selection, WikiHow Text-Image / Image-Text Temporal Ordering)
    10. Miscellaneous (e.g., Multimodal Factual Checking, Text Legibility, Text Type Classification, Image Captioning, Visual Text Extraction, Visual Dialogue, Disaster Type Classification)

    The tasks are sourced from 21 existing datasets (yielding 34 original tasks) and augmented with 28 newly derived tasks constructed by reformulating input and output structures. Newly created tasks contain between 5,000 and 5,000,000 instances each. Each task in MultiInstruct is equipped with 5 expert-written natural language instructions produced via an iterative review and refinement annotation process. The dataset is partitioned into 53 training tasks (designed to align with general pre-training competencies) and 9 challenging, non-overlapping evaluation tasks reserved for zero-shot testing.

  2. Knowl 2 — Unified Multimodal Sequence-to-Sequence Instruction Formatting

    model/method

    All multimodal tasks in MultiInstruct are represented in a unified sequence-to-sequence generation framework where text, images, bounding boxes, and task instructions share a single token vocabulary:

    • Text Representation: Input and output text sequences are tokenized using Byte-Pair Encoding (BPE).
    • Image Representation: Images are quantized into discrete visual tokens using a VQ-GAN codebook. When an input instance lacks an image, a solid black image is supplied as the visual input.
    • Bounding Box and Region Representation: Coordinate bounding boxes [xmin⁡,ymin⁡,xmax⁡,ymax⁡][x_{\min}, y_{\min}, x_{\max}, y_{\max}] are mapped to discrete location tokens ⟨bin_u⟩\langle\text{bin\_}u\rangle by dividing image dimensions into 1,000 bins (ranging from ⟨bin_0⟩\langle\text{bin\_0}\rangle to ⟨bin_999⟩\langle\text{bin\_999}\rangle). For example, a bounding box is represented as four tokens: "<bin_242> <bin_180> <bin_736> <bin_475>".
    • Task Instructions and Placeholders: Instructions are written as natural language templates containing arbitrary slots such as <TEXT>, <REGION>, <QUESTION>, <HISTORY>, <STEP>, <TASK>, and <OPTION>, which are instantiated with instance-specific data.
    • Candidate Options Representation: For classification tasks, the <OPTION> field introduces two special delimiter tokens: "[Options]" marking the start of candidate choices, and "||||" separating individual options. The model generates one of the specified candidate options autoregressively.
  3. Knowl 3 — Instruction Sensitivity Metric

    equation

    Sensitivity measures the degree of variance in a model's performance on a fixed task when prompted with different, semantically equivalent human-written instruction templates. Lower sensitivity indicates higher robustness to instruction phrasing.

    Given a set of evaluation tasks TT, where each task t∈Tt \in T is associated with a dataset Dt={(It,xjt,yjt)}j=1ND^t = \{(I^t, x_j^t, y_j^t)\}_{j=1}^{N} of inputs xjtx_j^t, ground truth outputs yjty_j^t, and a set of human-authored instructions It={i1,i2,…,iK}I^t = \{i_1, i_2, \dots, i_K\} (where K=5K=5), the sensitivity of an instruction-tuned model fθf_\theta under evaluation metric L\mathcal{L} (e.g., accuracy or ROUGE-L) is defined as:

    Sensitivity=Et∈T[σi∈It(E(x,y)∈Dt[L(fθ(i,x),y)])μi∈It(E(x,y)∈Dt[L(fθ(i,x),y)])]\text{Sensitivity} = \mathbb{E}_{t \in T} \left[ \frac{\sigma_{i \in I^t} \left( \mathbb{E}_{(x, y) \in D^t} [\mathcal{L}(f_\theta(i, x), y)] \right)}{\mu_{i \in I^t} \left( \mathbb{E}_{(x, y) \in D^t} [\mathcal{L}(f_\theta(i, x), y)] \right)} \right]

    where:

    • μi∈It[E(x,y)∈Dt[L(fθ(i,x),y)]]\mu_{i \in I^t} \left[ \mathbb{E}_{(x, y) \in D^t} [\mathcal{L}(f_\theta(i, x), y)] \right] is the mean performance of model fθf_\theta across all KK instructions for task tt.
    • σi∈It[E(x,y)∈Dt[L(fθ(i,x),y)]]\sigma_{i \in I^t} \left[ \mathbb{E}_{(x, y) \in D^t} [\mathcal{L}(f_\theta(i, x), y)] \right] is the sample standard deviation of model performance across all KK instructions for task tt.
    • Et∈T[⋅]\mathbb{E}_{t \in T}[\cdot] computes the unweighted average across all evaluation tasks in TT.
  4. Knowl 4 — Transfer Learning Strategies from Natural Instructions

    model/method

    To transfer instruction-following capabilities from large-scale text-only data to multimodal models, two transfer learning strategies incorporate 832 English tasks from the Natural Instructions benchmark into OFA:

    1. Mixed Instruction Tuning (OFAMixedInstruct\text{OFA}_{\text{MixedInstruct}}): Multimodal training instances from MultiInstruct and text-only instances from Natural Instructions are pooled together and randomly shuffled for joint instruction tuning. For each MultiInstruct instance, an instruction template is randomly sampled from its 5 available expert templates per training batch, while Natural Instructions instances use their single human-authored task definition.

    2. Sequential Instruction Tuning (OFASeqInstruct\text{OFA}_{\text{SeqInstruct}}): A two-stage training scheme inspired by pre-finetuning. In the first stage, OFA is fine-tuned exclusively on all English tasks in Natural Instructions to instill general text-based instruction understanding. In the second stage, the model is further fine-tuned on MultiInstruct to adapt instruction-following behaviors to visual and multimodal inputs.

  5. Knowl 5 — Zero-Shot Performance on Unseen Multimodal Tasks

    data/table

    Zero-shot evaluation is performed on 9 unseen multimodal tasks: Commonsense VQA (VCR), Visual Entailment (SNLI-VE), Visual Spatial Reasoning (VSR), Natural Language for Visual Reasoning (NLVR), Text VQA, Grounded VQA (Visual7W), Visual Text Extraction (Hateful Memes), Visual Dialogue, and Disaster Type Classification (MEDIC). For each task, models are tested over 5 independent instructions, reporting Max and Average ±\pm Standard Deviation.

    Task Metric OFA OFATaskName_{\text{TaskName}} OFAMultiInstruct_{\text{MultiInstruct}} OFAMixedInstruct_{\text{MixedInstruct}} OFASeqInstruct_{\text{SeqInstruct}}
    Commonsense VQA ROUGE-L 14.97 ±\pm 4.30 48.99 50.60 ±\pm 1.12 49.34 ±\pm 1.04 50.07 ±\pm 1.07
    Commonsense VQA ACC 0.40 ±\pm 0.29 29.01 31.17 ±\pm 1.59 30.27 ±\pm 0.94 31.23 ±\pm 1.09
    Visual Entailment ACC 41.86 ±\pm 10.99 55.70 55.06 ±\pm 0.76 53.74 ±\pm 0.97 52.98 ±\pm 0.56
    Visual Spatial Reasoning ACC 35.29 ±\pm 22.21 53.76 53.90 ±\pm 1.38 52.61 ±\pm 1.64 53.11 ±\pm 1.45
    NLVR ACC 52.10 ±\pm 3.35 55.35 56.18 ±\pm 0.95 55.96 ±\pm 0.48 56.63 ±\pm 0.66
    Text VQA ROUGE-L 9.30 ±\pm 5.42 23.80 26.46 ±\pm 0.83 23.67 ±\pm 0.47 26.67 ±\pm 0.47
    Grounded VQA ACC 0.00 ±\pm 0.01 0.00 47.22 ±\pm 23.08 54.99 ±\pm 18.16 54.46 ±\pm 15.96
    Visual Text Extraction ROUGE-L 17.62 ±\pm 16.82 36.30 62.43 ±\pm 11.56 46.56 ±\pm 14.92 60.62 ±\pm 12.31
    Visual Dialogue ROUGE-L 28.71 ±\pm 9.81 25.18 32.91 ±\pm 7.59 38.02 ±\pm 5.25 35.10 ±\pm 6.92
    Disaster Classification ACC 9.64 ±\pm 4.34 62.65 56.00 ±\pm 12.96 64.31 ±\pm 2.39 57.89 ±\pm 9.51

    Fine-tuning on MultiInstruct consistently improves zero-shot performance over base OFA and task-name conditioning (OFATaskName\text{OFA}_{\text{TaskName}}). On Grounded VQA, which requires generating quantized spatial tokens, base OFA and OFATaskName\text{OFA}_{\text{TaskName}} achieve 0.00% accuracy because they fail to follow the request for bounding box outputs, whereas instruction-tuned models reach 47.22% to 54.99%.

  6. Knowl 6 — Cross-Modal Attention Degradation under Text-Only Instruction Tuning

    empirical result

    When a vision-language model (OFA) is fine-tuned solely on text-only instructions (OFANaturalInstruct\text{OFA}_{\text{NaturalInstruct}}), its zero-shot multimodal performance degrades severely across nearly all vision-language tasks (e.g., dropping to 1.24 ROUGE-L on Visual Text Extraction and 0.00% on Grounded VQA).

    An analysis of the 12 encoder self-attention layers (16 attention heads each) reveals the mechanism causing this failure: computing the attention weights assigned by text query tokens to image key tokens shows that OFANaturalInstruct\text{OFA}_{\text{NaturalInstruct}} exhibits a sharp drop in text-to-image attention across all layers compared to base OFA and multimodal-tuned models. The decline is most acute in the initial two encoder layers, where text-to-image attention falls to approximately 0.050.05 (compared to 0.150.15--0.200.20 in base OFA and OFAMultiInstruct\text{OFA}_{\text{MultiInstruct}}). Tuning exclusively on text causes the encoder to ignore visual tokens. In contrast, joint mixed tuning (OFAMixedInstruct\text{OFA}_{\text{MixedInstruct}}) and sequential tuning (OFASeqInstruct\text{OFA}_{\text{SeqInstruct}}) preserve healthy cross-modal attention scores.

  7. Knowl 7 — Effect of Instruction Diversity on Zero-Shot Generalization and Sensitivity

    empirical result

    Varying the number of unique human-authored instructions provided per task during multimodal instruction tuning directly impacts downstream zero-shot accuracy and prompt robustness:

    • 1 Instruction per Task: Training OFAMultiInstruct\text{OFA}_{\text{MultiInstruct}} with a single fixed instruction template per task yields an aggregated zero-shot performance of 42.81 across unseen evaluation tasks and an instruction sensitivity of 24.62.
    • 5 Instructions per Task: Training OFAMultiInstruct\text{OFA}_{\text{MultiInstruct}} by randomly sampling among 5 diverse expert-authored instruction templates per task increases aggregated zero-shot performance to 47.82 (+5.01+5.01 points) and reduces instruction sensitivity to 10.45 (a 57.6%57.6\% reduction in variation).

    Exposing the model to linguistic diversity in instruction phrasing during training forces it to ground semantics rather than memorize specific prompt surface forms, yielding higher generalization and lower sensitivity to prompt variations.

  8. Knowl 8 — Impact of Task Cluster Scaling on Generalization and Sensitivity

    empirical result

    Incrementally scaling the number and diversity of task clusters during instruction tuning steadily improves model generalization while reducing instruction sensitivity. Grouping MultiInstruct and Natural Instructions tasks into cumulative clusters demonstrates this scaling trajectory:

    1. + Image Understanding (16 tasks): Initial aggregated performance ≈37.0\approx 37.0, sensitivity ≈36.5\approx 36.5.
    2. + Grounding (16 tasks; total 32 tasks): Performance increases to ≈43.0\approx 43.0, sensitivity drops to ≈23.0\approx 23.0.
    3. + MISC, ITM (14 tasks; total 46 tasks): Performance reaches ≈46.5\approx 46.5, sensitivity drops to ≈18.0\approx 18.0.
    4. + Visual Relationship (6 tasks; total 52 tasks): Performance reaches ≈47.5\approx 47.5, sensitivity drops to ≈15.0\approx 15.0.
    5. + Region Understanding (6 tasks; total 58 tasks): Performance reaches ≈48.0\approx 48.0, sensitivity drops to ≈14.0\approx 14.0.
    6. + NLP Tasks (832 tasks from Natural Instructions; total 890 tasks): Average aggregated performance reaches ≈50.0\approx 50.0 (Max ≈51.5\approx 51.5) and sensitivity drops to ≈10.27\approx 10.27.

    Both mean and maximum aggregated zero-shot performance increase monotonically with the number of instruction task clusters, accompanied by a continuous reduction in sensitivity.

  9. Knowl 9 — Instruction Sensitivity Reduction across Training Strategies

    empirical result

    Evaluating model sensitivity across all 9 unseen multimodal test tasks shows that instruction tuning substantially stabilizes model outputs against prompt perturbations:

    • Base OFA: Sensitivity = 40.58
    • OFAMultiInstruct\text{OFA}_{\text{MultiInstruct}}: Sensitivity = 13.84 (65.9% reduction relative to base OFA)
    • OFASeqInstruct\text{OFA}_{\text{SeqInstruct}}: Sensitivity = 10.45 (74.2% reduction relative to base OFA)
    • OFAMixedInstruct\text{OFA}_{\text{MixedInstruct}}: Sensitivity = 10.27 (74.7% reduction relative to base OFA)

    Instruction tuning directly mitigates language model brittleness to prompt wording. Furthermore, augmenting multimodal instruction tuning with large-scale text-only instruction datasets further reduces the sensitivity standard deviation across 6 out of 9 unseen multimodal tasks.

  10. Knowl 10 — Cross-Modal Transfer to Unseen NLP Tasks

    empirical result

    Evaluating multimodal instruction-tuned models on 20 text-only NLP test tasks from Natural Instructions (evaluated via ROUGE-L using the task 'Definition' as prompt) demonstrates cross-modal transfer:

    Model ROUGE-L
    OFA (Pre-trained Base) 2.25
    OFAMultiInstruct_{\text{MultiInstruct}} 12.18
    Transfer Learning from Natural Instructions
    OFANaturalInstruct_{\text{NaturalInstruct}} 43.61
    OFAMixedInstruct_{\text{MixedInstruct}} 43.32
    OFASeqInstruct_{\text{SeqInstruct}} 30.79

    Multimodal instruction tuning alone (OFAMultiInstruct\text{OFA}_{\text{MultiInstruct}}) improves text-only zero-shot ROUGE-L from 2.252.25 to 12.1812.18, demonstrating positive cross-modal transfer from multimodal instructions to pure text tasks. Mixed instruction tuning (OFAMixedInstruct\text{OFA}_{\text{MixedInstruct}}) matches the performance of the pure text-tuned model (43.3243.32 vs. 43.6143.61) while avoiding the catastrophic forgetting observed in sequential tuning (30.7930.79 ROUGE-L).

  11. Knowl 11 — Limitations of MultiInstruct and Evaluation Protocol

    limitation

    The MultiInstruct dataset and experimental methodology have several explicitly stated limitations:

    1. Language Scope: The dataset is restricted to English-language tasks and instructions, lacking multilingual coverage.
    2. Modality Constraints: The benchmark is confined to vision-and-language tasks and does not incorporate other sensory modalities such as audio or video.
    3. Scale of Tasks and Instructions: With 62 tasks and 5 instructions per task, the prompt diversity remains constrained compared to massive text-only instruction datasets.
    4. Model Diversity: Experiments rely solely on OFA-large (472M parameters) as the base pretrained model, without evaluating models across varied parameter scales or alternative architectures.
    5. Metric Scope: The sensitivity metric only measures intra-task variance across alternative instructions for the same task, without measuring inter-task instruction discrimination capabilities.

Coverage note — Omitted the full enumeration of all 53 training tasks and individual dataset bibliographies from the appendix tables, as their categorical distribution, derivation methodology, and representative task structures are fully preserved in the primary dataset and result knowls.

References

  1. 1.Armen Aghajanyan, Anchit Gupta, Akshat Shrivastava, Xilun Chen, Luke Zettlemoyer, and Sonal Gupta. 2021. Muppet: Massive multi-task representations with pre-finetuning. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 5799–5811, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  2. 2.Firoj Alam, Tanvirul Alam, Md Hasan, Abul Hasnat, Muhammad Imran, Ferda Ofli, et al. 2022. Medic: a multi-task learning dataset for disaster image classification. Neural Computing and Applications, pages 1–24.
  3. 3.Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katie Millican, Malcolm Reynolds, et al. 2022. Flamingo: a visual language model for few-shot learning. arXiv preprint arXiv:2204.14198.
  4. 4.Hangbo Bao, Li Dong, Songhao Piao, and Furu Wei. 2022. Beit: Bert pre-training of image transformers. In ICLR 2022.
  5. 5.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901.
  6. 6.Tai-Yin Chiu, Yinan Zhao, and Danna Gurari. 2020. Assessing image quality issues for real-world problems. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3646–3656.
  7. 7.Jaemin Cho, Jie Lei, Hao Tan, and Mohit Bansal. 2021. Unifying vision-and-language tasks via text generation. In International Conference on Machine Learning, pages 1931–1942. PMLR.
  8. 8.Abhishek Das, Satwik Kottur, Khushi Gupta, Avi Singh, Deshraj Yadav, José MF Moura, Devi Parikh, and Dhruv Batra. 2017. Visual dialog. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 326–335.
  9. 9.Patrick Esser, Robin Rombach, and Bjorn Ommer. 2021. Taming transformers for high-resolution image synthesis. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 12873–12883.
  10. 10.Jort F. Gemmeke, Daniel P. W. Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R. Channing Moore, Manoj Plakal, and Marvin Ritter. 2017. Audio set: An ontology and human-labeled dataset for audio events. In 2017 IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP 2017, New Orleans, LA, USA, March 5-9, 2017, pages 776–780. IEEE.
  11. 11.Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh. 2017. Making the v in vqa matter: Elevating the role of image understanding in visual question answering. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 6904–6913.
  12. 12.Prakhar Gupta, Cathy Jiao, Yi-Ting Yeh, Shikib Mehri, Maxine Eskenazi, and Jeffrey P. Bigham. 2022. Improving zero and few-shot generalization in dialogue through instruction tuning.
  13. 13.Xu Han, Weilin Zhao, Ning Ding, Zhiyuan Liu, and Maosong Sun. 2022. Ptr: Prompt tuning with rules for text classification. AI Open.
  14. 14.Drew A Hudson and Christopher D Manning. 2019. Gqa: A new dataset for real-world visual reasoning and compositional question answering. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 6700–6709.
  15. 15.Catalin Ionescu, Dragos Papava, Vlad Olaru, and Cristian Sminchisescu. 2014. Human3.6m: Large scale datasets and predictive methods for 3d human sensing in natural environments. IEEE Trans. Pattern Anal. Mach. Intell., 36(7):1325–1339.
  16. 16.Kushal Kafle and Christopher Kanan. 2017. An analysis of visual question answering algorithms. In Proceedings of the IEEE international conference on computer vision, pages 1965–1973.
  17. 17.Douwe Kiela, Hamed Firooz, Aravind Mohan, Vedanuj Goswami, Amanpreet Singh, Pratik Ringshia, and Davide Testuggine. 2020. The hateful memes challenge: Detecting hate speech in multimodal memes. Advances in Neural Information Processing Systems, 33:2611–2624.
  18. 18.Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A Shamma, et al. 2017. Visual genome: Connecting language and vision using crowdsourced dense image annotations. International journal of computer vision, 123(1):32–73.
  19. 19.Hao Li, Jinguo Zhu, Xiaohu Jiang, Xizhou Zhu, Hongsheng Li, Chun Yuan, Xiaohua Wang, Yu Qiao, Xiaogang Wang, Wenhai Wang, and Jifeng Dai. 2022a. Uni-perceiver v2: A generalist model for large-scale vision and vision-language tasks. CoRR, abs/2211.09808.
  20. 20.Linjie Li, Zhe Gan, Kevin Lin, Chung-Ching Lin, Zicheng Liu, Ce Liu, and Lijuan Wang. 2022b. Lavender: Unifying video-language understanding as masked language modeling. arXiv preprint arXiv:2206.07160.
  21. 21.Xiang Lisa Li and Percy Liang. 2021. Prefix-tuning: Optimizing continuous prompts for generation. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, ACL/IJCNLP 2021, (Volume 1: Long Papers), Virtual Event, August 1-6, 2021, pages 4582–4597. Association for Computational Linguistics.
  22. 22.Chin-Yew Lin. 2004. Rouge: A package for automatic evaluation of summaries. In Text summarization branches out, pages 74–81.
  23. 23.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. 2014. Microsoft coco: Common objects in context. In European conference on computer vision, pages 740–755. Springer.
  24. 24.Fangyu Liu, Guy Emerson, and Nigel Collier. 2022a. Visual spatial reasoning. arXiv preprint arXiv:2205.00363.
  25. 25.Haokun Liu, Derek Tam, Mohammed Muqeeth, Jay Mohta, Tenghao Huang, Mohit Bansal, and Colin Raffel. 2022b. Few-shot parameter-efficient fine-tuning is better and cheaper than in-context learning. CoRR, abs/2205.05638.
  26. 26.Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig. 2021. Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing. arXiv preprint arXiv:2107.13586.
  27. 27.Jiasen Lu, Christopher Clark, Rowan Zellers, Roozbeh Mottaghi, and Aniruddha Kembhavi. 2022. Unified-io: A unified model for vision, language, and multimodal tasks.
  28. 28.Kenneth Marino, Mohammad Rastegari, Ali Farhadi, and Roozbeh Mottaghi. 2019. Ok-vqa: A visual question answering benchmark requiring external knowledge. In Proceedings of the IEEE/cvf conference on computer vision and pattern recognition, pages 3195–3204.
  29. 29.Sewon Min, Mike Lewis, Luke Zettlemoyer, and Hannaneh Hajishirzi. 2021. Metaicl: Learning to learn in context. arXiv preprint arXiv:2110.15943.
  30. 30.Swaroop Mishra, Daniel Khashabi, Chitta Baral, and Hannaneh Hajishirzi. 2022. Cross-task generalization via natural language crowdsourcing instructions. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 3470–3487, Dublin, Ireland. Association for Computational Linguistics.
  31. 31.Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur. 2015. Librispeech: An ASR corpus based on public domain audio books. In 2015 IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP 2015, South Brisbane, Queensland, Australia, April 19-24, 2015, pages 5206–5210. IEEE.
  32. 32.Khoi Pham, Kushal Kafle, Zhe Lin, Zhihong Ding, Scott Cohen, Quan Tran, and Abhinav Shrivastava. 2021. Learning to predict visual attributes in the wild. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 13018–13028.
  33. 33.Victor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Arun Raja, Manan Dey, M Saiful Bari, Canwen Xu, Urmish Thakker, Shanya Sharma Sharma, Eliza Szczechla, Taewoon Kim, Gunjan Chhablani, Nihal V. Nayak, Debajyoti Datta, Jonathan Chang, Mike Tian-Jian Jiang, Han Wang, Matteo Manica, Sheng Shen, Zheng Xin Yong, Harshit Pandey, Rachel Bawden, Thomas Wang, Trishala Neeraj, Jos Rozen, Abheesht Sharma, Andrea Santilli, Thibault Févry, Jason Alan Fries, Ryan Teehan, Teven Le Scao, Stella Biderman, Leo Gao, Thomas Wolf, and Alexander M. Rush. 2022. Multitask prompted training enables zero-shot task generalization. In The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 25-29, 2022. OpenReview.net.
  34. 34.Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016. Neural machine translation of rare words with subword units. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1715–1725.
  35. 35.Amanpreet Singh, Ronghang Hu, Vedanuj Goswami, Guillaume Couairon, Wojciech Galuba, Marcus Rohrbach, and Douwe Kiela. 2022. Flava: A foundational language and vision alignment model. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 15638–15650.
  36. 36.Amanpreet Singh, Vivek Natarajan, Meet Shah, Yu Jiang, Xinlei Chen, Dhruv Batra, Devi Parikh, and Marcus Rohrbach. 2019. Towards vqa models that can read. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 8317–8326.
  37. 37.Khurram Soomro, Amir Roshan Zamir, and Mubarak Shah. 2012. UCF101: A dataset of 101 human actions classes from videos in the wild. CoRR, abs/1212.0402.
  38. 38.Alane Suhr, Mike Lewis, James Yeh, and Yoav Artzi. 2017. A corpus of natural language for visual reasoning. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 217–223.
  39. 39.Hao Tan and Mohit Bansal. 2019. Lxmert: Learning cross-modality encoder representations from transformers. arXiv preprint arXiv:1908.07490.
  40. 40.Andreas Veit, Tomas Matera, Lukas Neumann, Jiri Matas, and Serge Belongie. 2016. Coco-text: Dataset and benchmark for text detection and recognition in natural images. arXiv preprint arXiv:1601.07140.
  41. 41.Peng Wang, An Yang, Rui Men, Junyang Lin, Shuai Bai, Zhikang Li, Jianxin Ma, Chang Zhou, Jingren Zhou, and Hongxia Yang. 2022a. Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework. arXiv preprint arXiv:2202.03052.
  42. 42.Sijia Wang, Mo Yu, and Lifu Huang. 2022b. The art of prompting: Event detection based on type specific prompts. arXiv preprint arXiv:2204.07241.
  43. 43.Wenhui Wang, Hangbo Bao, Li Dong, Johan Bjorck, Zhiliang Peng, Qiang Liu, Kriti Aggarwal, Owais Khan Mohammed, Saksham Singhal, Subhojit Som, and Furu Wei. 2022c. Image as a foreign language: Beit pretraining for all vision and vision-language tasks. CoRR, abs/2208.10442.
  44. 44.Yizhong Wang, Swaroop Mishra, Pegah Alipoormolabashi, Yeganeh Kordi, Amirreza Mirzaei, Anjana Arunkumar, Arjun Ashok, Arut Selvan Dhanasekaran, Atharva Naik, David Stap, Eshaan Pathak, Giannis Karamanolakis, Haizhi Gary Lai, Ishan Purohit, Ishani Mondal, Jacob Anderson, Kirby Kuznia, Krima Doshi, Maitreya Patel, Kuntal Kumar Pal, Mehrad Moradshahi, Mihir Parmar, Mirali Purohit, Neeraj Varshney, Phani Rohitha Kaza, Pulkit Verma, Ravsehaj Singh Puri, Rushang Karia, Shailaja Keyur Sampat, Savan Doshi, Siddhartha Mishra, Sujan Reddy, Sumanta Patro, Tanay Dixit, Xudong Shen, Chitta Baral, Yejin Choi, Noah A. Smith, Hannaneh Hajishirzi, and Daniel Khashabi. 2022d. Super-naturalinstructions: Generalization via declarative instructions on 1600+ nlp tasks.
  45. 45.Albert Webson and Ellie Pavlick. 2022. Do prompt-based models really understand the meaning of their prompts? In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2300–2344, Seattle, United States. Association for Computational Linguistics.
  46. 46.Jason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M. Dai, and Quoc V. Le. 2021. Finetuned language models are zero-shot learners. CoRR, abs/2109.01652.
  47. 47.Ning Xie, Farley Lai, Derek Doran, and Asim Kadav. 2019. Visual entailment: A novel task for fine-grained image understanding. arXiv preprint arXiv:1901.06706.
  48. 48.Sang Michael Xie, Aditi Raghunathan, Percy Liang, and Tengyu Ma. 2021. An explanation of in-context learning as implicit bayesian inference. CoRR, abs/2111.02080.
  49. 49.Barry Menglong Yao, Aditya Shah, Lichao Sun, Jin-Hee Cho, and Lifu Huang. 2022. End-to-end multimodal fact-checking and explanation generation: A challenging dataset and models. arXiv preprint arXiv:2205.12487.
  50. 50.Chenyu You, Nuo Chen, Fenglin Liu, Shen Ge, Xian Wu, and Yuexian Zou. 2022. End-to-end spoken conversational question answering: Task, dataset and model. In Findings of the Association for Computational Linguistics: NAACL 2022, Seattle, WA, United States, July 10-15, 2022, pages 1219–1232. Association for Computational Linguistics.
  51. 51.Licheng Yu, Patrick Poirson, Shan Yang, Alexander C Berg, and Tamara L Berg. 2016. Modeling context in referring expressions. In European Conference on Computer Vision, pages 69–85. Springer.
  52. 52.Rowan Zellers, Yonatan Bisk, Ali Farhadi, and Yejin Choi. 2019. From recognition to cognition: Visual commonsense reasoning. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 6720–6731.
  53. 53.Yuke Zhu, Oliver Groth, Michael Bernstein, and Li Fei-Fei. 2016. Visual7w: Grounded question answering in images. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 4995–5004.

Citation

MLA
Xu, Z., et al. “MultiInstruct: Improving Multi-Modal Zero-Shot Learning via Instruction Tuning”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2023, pp. 11445–65, https://doi.org/10.18653/v1/2023.acl-long.641.
APA
Xu, Z., Shen, Y., & Huang, L. (2023). MultiInstruct: Improving Multi-Modal Zero-Shot Learning via Instruction Tuning. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 11445–11465. https://doi.org/10.18653/v1/2023.acl-long.641
Chicago
Xu, Z., Y. Shen, and L. Huang. 2023. “MultiInstruct: Improving Multi-Modal Zero-Shot Learning via Instruction Tuning”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 11445–65. https://doi.org/10.18653/v1/2023.acl-long.641.
Harvard
Xu, Z., Shen, Y. and Huang, L. (2023) “MultiInstruct: Improving Multi-Modal Zero-Shot Learning via Instruction Tuning”, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 11445–11465. Available at: https://doi.org/10.18653/v1/2023.acl-long.641.
Vancouver
1. Xu Z, Shen Y, Huang L (2023) MultiInstruct: Improving Multi-Modal Zero-Shot Learning via Instruction Tuning. In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 11445–11465

BibTeX

@inproceedings{xu-etal-2023-multiinstruct,
    title = "{M}ulti{I}nstruct: Improving Multi-Modal Zero-Shot Learning via Instruction Tuning",
    author = "Xu, Zhiyang  and
      Shen, Ying  and
      Huang, Lifu",
    editor = "Rogers, Anna  and
      Boyd-Graber, Jordan  and
      Okazaki, Naoaki",
    booktitle = "Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2023",
    address = "Toronto, Canada",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.acl-long.641/",
    doi = "10.18653/v1/2023.acl-long.641",
    pages = "11445--11465"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/