OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework

Peng WangAn YangRui MenJunyang LinShuai BaiZhikang LiJianxin MaChang ZhouJingren ZhouHongxia Yang

article2022ICML1,052 citations

Introduces OFA, an instruction-driven sequence-to-sequence framework that unifies diverse vision and language tasks—from image generation to visual grounding—into a single architecture without task-specific layers, achieving state-of-the-art performance using only 20 million pretraining image-text pairs.

Listen

Current artificial intelligence systems often require separate, customized model architectures and task-specific components for different types of data, such as images, text, and combined visual-textual inputs. This fragmentation increases system complexity, inflates training and maintenance costs, and limits the ability of models to generalize across varied, real-world tasks. The article addresses this challenge by introducing OFA (One For All), a unified framework designed to process multimodal and unimodal tasks—including text generation, visual grounding, image classification, and image generation—within a single, standardized architecture without adding custom layers for downstream applications.

The core objective of the article is to demonstrate that an omnipotent model can achieve task-agnostic and modality-agnostic capabilities through a sequence-to-sequence learning structure. The authors evaluate this approach using an encoder-decoder Transformer backbone pretrained on a comparatively compact dataset of 20 million publicly available image-text pairs, accompanied by visual and text-only corpora. By converting text, image patches, and bounding box coordinates into a unified token vocabulary and conditioning tasks via natural language instructions, the framework executes diverse tasks using the same compute engine.

The findings show that OFA outperforms or matches existing state-of-the-art models across several benchmarks. In cross-modal benchmarks, OFA achieved top results, including an 82.0 accuracy score on Visual Question Answering (test-std) and a leading 154.9 CIDEr score on MSCOCO image captioning. For image generation, OFA achieved an improved Fréchet Inception Distance of 10.5 while utilizing a significantly smaller candidate sample size (24 samples) compared to competing baselines. On unimodal benchmarks, the framework achieved performance comparable to leading domain-specific systems, matching advanced language models on standard language understanding benchmarks, establishing a state-of-the-art score on Gigaword text summarization, and attaining an 85.6% top-1 accuracy on ImageNet-1K image classification. Furthermore, the model exhibited solid zero-shot transfer capabilities to unseen tasks, such as grounded question answering, and out-of-domain visual data.

These results carry significant implications for artificial intelligence development and deployment. By proving that a unified sequence-to-sequence structure can outperform specialized models without requiring massive proprietary datasets—such as those scaling to nearly two billion pairs—OFA presents an efficient path toward lowering computational overhead, simplifying model lifecycle management, and minimizing engineering risk. Organizations can leverage a single foundation model across diverse operational needs rather than building and maintaining isolated models for each application.

Based on these outcomes, engineering and research teams should consider adopting unified sequence-to-sequence architectures to consolidate their artificial intelligence infrastructure. However, stakeholders should note specific operational limitations: the model exhibits high sensitivity to natural language prompt formulation, and its zero-shot capabilities remain constrained on sentence-pair classification tasks due to pretraining data boundaries. Further work is recommended to automate prompt optimization and systematically investigate broader pretraining data distributions to strengthen zero-shot robustness before deploying the system in critical, unsupervised production environments.

Cover for OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework

Abstract

In this work, we pursue a unified paradigm for multimodal pretraining to break the scaffolds of complex task/modality-specific customization. We propose OFA, a Task-Agnostic and Modality-Agnostic framework that supports Task Comprehensiveness. OFA unifies a diverse set of cross-modal and unimodal tasks, including image generation, visual grounding, image captioning, image classification, language modeling, etc., in a simple sequence-to-sequence learning framework. OFA follows the instruction-based learning in both pretraining and finetuning stages, requiring no extra task-specific layers for downstream tasks. In comparison with the recent state-of-the-art vision & language models that rely on extremely large cross-modal datasets, OFA is pretrained on only 20M publicly available image-text pairs. Despite its simplicity and relatively small-scale training data, OFA achieves new SOTAs in a series of cross-modal tasks while attaining highly competitive performances on uni-modal tasks. Our further analysis indicates that OFA can also effectively transfer to unseen tasks and unseen domains. Our code and models are publicly available at this https URL.

Citation

MLA
Wang, P., et al. “OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework”. arXiv, 2022, http://arxiv.org/abs/2202.03052v2.
APA
Wang, P., Yang, A., Men, R., Lin, J., Bai, S., Li, Z., Ma, J., Zhou, C., Zhou, J., & Yang, H. (2022). OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework. arXiv. http://arxiv.org/abs/2202.03052v2
Chicago
Wang, P., A. Yang, R. Men, et al. 2022. “OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework”. arXiv. http://arxiv.org/abs/2202.03052v2.
Harvard
Wang, P. et al. (2022) “OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2202.03052v2.
Vancouver
1. Wang P, Yang A, Men R, Lin J, Bai S, Li Z, Ma J, Zhou C, Zhou J, Yang H (2022) OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework. arXiv

BibTeX

@article{wang2022ofa,
  title = {OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework},
  author = {Wang, Peng and Yang, An and Men, Rui and Lin, Junyang and Bai, Shuai and Li, Zhikang and Ma, Jianxin and Zhou, Chang and Zhou, Jingren and Yang, Hongxia},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2202.03052v2},
  eprint = {2202.03052}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/