Built independently by an author, for readers. Read the story and support ChapterPal

keyword

unified architecture

A unified architecture is a single neural network framework designed to process diverse data modalities and perform multiple distinct tasks within a shared model structure rather than relying on specialized, task-specific or modality-specific components. By converting varied data types such as text, images, and audio into a shared representation or sequence format, this design enables end-to-end learning across unimodal and cross-modal applications ranging from classification and reasoning to generative tasks. This standardized formulation simplifies model training and fine-tuning, facilitates broad knowledge transfer across domains, and allows a single system to generalize effectively to new or unseen tasks without requiring structural modifications.

1 item

OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework

OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework

Peng Wang, An Yang, Rui Men, Junyang Lin, Shuai Bai, Zhikang Li, Jianxin Ma, Chang Zhou, Jingren Zhou, Hongxia Yang

OrganizationsAlibaba Group

Why you should read this

Introduces OFA, an instruction-driven sequence-to-sequence framework that unifies diverse vision and language tasks—from image generation to visual grounding—into a single architecture without task-specific layers, achieving state-of-the-art performance using only 20 million pretraining image-text pairs.

In this work, we pursue a unified paradigm for multimodal pretraining to break the scaffolds of complex task/modality-specific customization. We propose OFA, a Task-Agnostic and Modality-Agnostic framework that supports Task Comprehensiveness. OFA unifies a diverse set of cross-modal and unimodal tasks, including image generation, visual grounding, image captioning, image classification, language modeling, etc., in a simple sequence-to-sequence learning framework. OFA follows the instruction-based learning in both pretraining and finetuning stages, requiring no extra task-specific layers for downstream tasks. In comparison with the recent state-of-the-art vision & language models that rely on extremely large cross-modal datasets, OFA is pretrained on only 20M publicly available image-text pairs. Despite its simplicity and relatively small-scale training data, OFA achieves new SOTAs in a series of cross-modal tasks while attaining highly competitive performances on uni-modal tasks. Our further analysis indicates that OFA can also effectively transfer to unseen tasks and unseen domains. Our code and models are publicly available at this https URL.

Added

2026-09-28