MultiInstruct: Improving Multi-Modal Zero-Shot Learning via Instruction Tuning
Zhiyang XuYing ShenLifu Huang
Introduces the first multimodal instruction tuning benchmark spanning 62 diverse tasks across 10 categories to boost zero-shot generalization and reduce sensitivity to prompt phrasing in vision-language models.
Instruction tuning has enabled large language models to generalize to unseen natural language tasks without task-specific training, but its potential for multimodal systems remains largely unexplored. The article addresses this gap by introducing MULTIINSTRUCT, the first benchmark dataset designed for multimodal instruction tuning. The primary objective is to demonstrate that fine-tuning vision-language models on diverse multimodal tasks guided by natural language instructions significantly improves their zero-shot generalization to unseen tasks, while reducing sensitivity to variations in prompt wording.
To evaluate this approach, the authors compiled 62 multimodal tasks spanning 10 broad categories derived from 21 open-source datasets, pairing each task with five expert-written instruction templates. All tasks were mapped into a unified sequence-to-sequence format where text, images, and visual coordinates share a single vocabulary. Using the pre-trained OFA model as a base, the study fine-tuned the system on 53 tasks and evaluated zero-shot performance across nine unseen multimodal tasks and 20 text-only tasks. The authors also explored transfer learning strategies using NATURAL INSTRUCTIONS, a large-scale repository of 832 English text-only tasks, and introduced a Sensitivity metric to measure model output stability across different prompt wordings.
The findings show that multimodal instruction tuning substantially improves zero-shot capabilities. For instance, on Grounded Visual Question Answering, the baseline model scored near zero because it failed to follow spatial instructions, whereas the instruction-tuned model achieved an average accuracy of 47.22%. Training on multiple diverse instructions reduced model sensitivity to prompt variations by more than half compared to the base model. Notably, tuning exclusively on text-only instructions degraded vision-language performance because the attention layers shifted focus away from visual tokens. However, combining text-only and multimodal tasks simultaneously achieved the strongest overall performance, preserving text reasoning while improving multimodal stability.
These results demonstrate that multimodal instruction tuning is a practical and robust strategy for building flexible, generalist artificial intelligence systems that do not require custom re-training for new use cases. For organizations deploying vision-language models, the authors recommend incorporating instruction diversity during training and adopting mixed multimodal-text datasets rather than sequential or single-modality tuning. Future work should address current limitations by expanding beyond English datasets, incorporating audio and video modalities, and testing larger foundational models.
- Paper: Finetuned Language Models Are Zero-Shot Learners, Jason Wei et al. (2022). This seminal work establishes natural language instruction tuning across diverse tasks to induce zero-shot generalization, providing the foundational conceptual paradigm that MultiInstruct adapts to vision-language models.
- Paper: Multitask Prompted Training Enables Zero-Shot Task Generalization, Victor Sanh et al. (2021). It introduces multitask prompted fine-tuning and prompt-sensitivity evaluations across diverse NLP datasets, establishing the prompt formatting and robustness evaluation strategies built upon by MultiInstruct.
- Paper: Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks, Jiasen Lu et al. (2022). It defines a unified sequence-to-sequence framework mapping text, images, and visual coordinates into a shared discrete vocabulary, which underpins the unified input-output representation utilized in MultiInstruct.
- Paper: BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation, Junnan Li et al. (2022). It details unified vision-language pre-training for multimodal understanding and generation, providing core architectures and data-bootstrapping methods assumed by multimodal instruction tuning.
- Paper: Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering, Pan Lu et al. (2022). It provides a key benchmark and methodology for multimodal multi-step reasoning that serves as foundational context for evaluating instruction-guided visual question answering.
- Paper: Visual Instruction Tuning, Haotian Liu et al. (2023). It advances multimodal instruction tuning by using machine-generated conversational data from GPT-4 to train general-purpose visual assistants.
- Paper: InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning, Wenliang Dai et al. (2023). It extends multimodal instruction tuning by introducing an instruction-aware Query Transformer on top of BLIP-2 to condition visual feature extraction directly on textual prompts.
- Paper: Improved Baselines with Visual Instruction Tuning, Haotian Liu et al. (2024). It refines visual instruction tuning through enhanced connector designs and targeted visual task scaling, systematically improving upon early instruction-tuned baselines.
- Paper: MMBench: Is Your Multi-modal Model an All-around Player?, Yuanzhan Liu et al. (2023). It establishes a fine-grained evaluation benchmark that addresses instruction-following sensitivity and robustly measures the diverse multimodal capabilities promoted by MultiInstruct.
- Paper: MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models, Chaoyou Fu et al. (2023). It provides a comprehensive, objective benchmark to evaluate the perceptual and cognitive capabilities of instruction-tuned multimodal large language models.
- Paper: MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities, Weihao Yu et al. (2023). It extends the evaluation of multimodal instruction followers to complex open-ended tasks requiring the integration of multiple core vision-language capabilities.
- Paper: Active Instruction Tuning: Improving Cross-Task Generalization by Training on Prompt Sensitive Tasks, Po-Nien Kung et al. (2023). It directly builds on the prompt-sensitivity challenges highlighted in MultiInstruct by actively selecting training tasks that exhibit high prompt uncertainty.
- Paper: Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond, Jinze Bai et al. (2023). It scales unified multimodal instruction tuning across diverse fine-grained vision tasks including spatial grounding, document reading, and multilingual dialogue.
- Paper: MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models, Deyao Zhu et al. (2024). It explores instruction tuning for vision-language models by connecting advanced large language models with visual encoders using aligned conversational instruction datasets.
- Paper: A survey on multimodal large language models, Shukang Yin et al. (2023). It synthesizes the broader landscape of multimodal large language models, surveying the architecture, training pipelines, and instruction-tuning benchmarks that evolved after MultiInstruct.
