Uni-Perceiver-MoE: Learning Sparse Generalist Models with Conditional MoEs
Jinguo ZhuXizhou ZhuWenhai WangXiaohua WangHongsheng LiXiaogang WangJifeng Dai
Proposes Conditional Mixture-of-Experts routing strategies for generalist models to resolve cross-task and cross-modality parameter interference, achieving state-of-the-art multi-task and zero-shot performance with minimal data and computational cost.
Artificial intelligence research has increasingly shifted toward unified generalist models that can execute multiple tasks across text, image, and video modalities using a single shared set of parameters. However, when a single model handles diverse tasks simultaneously, it often experiences performance degradation compared to specialized, single-task systems. This issue, known as task interference, arises because the optimization gradients required for distinct objectives frequently point in conflicting directions, leading to suboptimal compromises within shared model parameters.
The main objective of the article is to identify and quantify the mechanics of task interference in generalist models and to demonstrate that a conditional routing framework can mitigate cross-task conflicts without sacrificing the model's ability to adapt or generalize to unseen downstream tasks.
To achieve this, the authors analyzed gradient behavior across various layers of a generalist model and evaluated several routing configurations within a Mixture-of-Experts (MoE) architecture. Unlike conventional approaches that introduce task-specific components or route computations solely based on token patterns, the authors designed Conditional Mixture-of-Experts (Conditional MoEs) and tested them across five condition levels: token, context, modality, task, and predefined token attributes. The experimental framework integrated these techniques into the Uni-Perceiver model family (across Tiny, Base, and Large configurations) using multi-modal datasets—including ImageNet, Books&Wiki, MSCOCO, Kinetics-400, Flickr30k, and GLUE—to evaluate both standard and zero-shot novel task performance.
The evaluation yielded several critical findings. First, gradient conflict is significantly more pronounced in deeper model layers than in shallow layers, confirming that deep representations suffer the most from shared-parameter multi-task interference. Second, routing experts using an 8-bit attribute embedding—which encodes input/target modalities, causality, and token source—achieved superior performance over token-, context-, and modality-level variants while avoiding the generalization bottlenecks of task-specific IDs. Third, when incorporating attribute-conditioned experts, the generalist model matched or outperformed task-specialized baselines across vision, language, and cross-modal tasks. Fourth, when tuned using lightweight prompts on only 1% of downstream data, the model achieved performance competitive with state-of-the-art benchmarks while consuming less than 5% of the training data and under 10% of the training cost. Finally, the model retained strong generalization, establishing robust zero-shot and few-shot results on unseen tasks such as video captioning and video-text retrieval.
These findings indicate that sparse parameter activation conditioned on task and modality attributes substantially improves multi-task learning efficiency. By reducing data dependency and computing requirements, organizations can train and deploy highly capable foundation models at a fraction of the cost, timeline, and memory overhead typical of massive dense systems. Furthermore, using reparameterization techniques for data-independent attribute routing eliminates the excessive memory and inter-device communication latency typically associated with standard Mixture-of-Experts architectures during deployment.
Based on these results, engineering teams and organizations deploying multi-modal models should consider adopting attribute-based conditional routing as an architectural standard to resolve task conflict. Decision-makers should leverage lightweight prompt tuning rather than complete model fine-tuning on large downstream datasets to minimize compute expenditure. Before broad enterprise deployment on giant architectures, teams should run targeted pilot programs to evaluate whether the scaling benefits observed in hundred-million-parameter models hold at multi-billion-parameter scales.
The study's primary limitation is that it evaluates models scaled up to several hundred million parameters, leaving the exact behavior of Conditional MoEs on billion-scale foundation models an open empirical question. Additionally, like all models pre-trained on large-scale web datasets, the system carries potential societal concerns regarding energy consumption during training and latent dataset biases. Nonetheless, for the evaluated scales and tasks, the findings provide high confidence that conditional sparse routing effectively resolves multi-task degradation.
- Paper: Modeling Task Relationships in Multi-task Learning with Multi-gate Mixture-of-Experts, Jiaqi Ma et al. (2018). Introduces task-conditioned gating across shared mixture-of-experts subnetworks to mitigate gradient conflicts in multi-task learning, laying foundational concepts directly built upon by Uni-Perceiver-MoE.
- Paper: Gradient Surgery for Multi-Task Learning, Tianhe Yu et al. (2020). Provides the seminal analysis and mitigation framework for conflicting task gradients in multi-task parameter optimization, motivating Uni-Perceiver-MoE's investigation into gradient interference.
- Paper: Scaling Vision with Sparse Mixture of Experts, Carlos Riquelme et al. (2021). Demonstrates the practical viability and routing mechanisms of sparse Mixture-of-Experts within Vision Transformer architectures, which Uni-Perceiver-MoE extends across modalities.
- Paper: Multi-Task Learning as Multi-Objective Optimization, Ozan Sener et al. (2018). Formulates multi-task representation sharing as multi-objective optimization to handle task gradient competition without capacity degradation.
- Paper: GradNorm: Gradient Normalization for Adaptive Loss Balancing in Deep Multitask Networks, Zhao Chen et al. (2017). Establishes gradient balancing techniques to resolve task dominance and interference across shared multi-task backbones.
- Paper: Adaptive Mixtures of Local Experts, Robert A. Jacobs et al. (1991). Presents the foundational mixture-of-experts paradigm with dynamic gating to eliminate task interference during joint learning.
- Paper: Learning Factored Representations in a Deep Mixture of Experts, David Eigen et al. (2014). Extends mixture-of-experts gating to deep neural network layers to learn factored and specialized subnetwork representations.
- Paper: End-To-End Multi-Task Learning With Attention, Shikun Liu et al. (2018). Explores architectural routing via attention mechanisms to separate task-specific and globally shared features in multi-task models.
- Paper: Multimodal Contrastive Learning with LIMoE: the Language-Image Mixture of Experts, Basil Mustafa et al. (2022). Applies sparse mixture-of-experts directly to multimodal text and image token routing within contrastive generalist architectures.
- Paper: Mixture-of-Experts with Expert Choice Routing, Yanqi Zhou et al. (2022). Introduces expert-choice routing strategies to resolve load imbalances and representation bottlenecks in sparse expert transformers.
- Paper: ST-MoE: Designing Stable and Transferable Sparse Expert Models, Barret Zoph et al. (2022). Develops router regularization and transfer recipes to stabilize and improve downstream task fine-tuning for sparse expert models.
- Paper: From Sparse to Soft Mixtures of Experts, Joan Puigcerver et al. (2024). Generalizes discrete token-to-expert routing to fully differentiable soft mixtures of experts for multi-task vision models.
- Paper: Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks, Jiasen Lu et al. (2022). Takes the unified generalist model concept further by homogenizing input and output representations across dozens of vision, language, and multimodal tasks into a single discrete sequence.
- Paper: PMR: Prototypical Modal Rebalance for Multimodal Learning, Yunfeng Fan et al. (2023). Investigates and resolves the problem of modality competition and dominance in multimodal representation learning.
- Paper: ReconBoost: Boosting Can Achieve Modality Reconcilement, Cong Hua et al. (2024). Proposes boosting-based reconcilement methods to prevent dominant modalities from suppressing weaker modalities during multimodal joint learning.
- Paper: VMT-Adapter: Parameter-Efficient Transfer Learning for Multi-Task Dense Scene Understanding, Yi Xin et al. (2024). Builds upon parameter-efficient multi-task adaptation by developing shared and task-specific adapters for dense multimodal vision tasks.
