Uni-Perceiver-MoE: Learning Sparse Generalist Models with Conditional MoEs

Jinguo ZhuXizhou ZhuWenhai WangXiaohua WangHongsheng LiXiaogang WangJifeng Dai

article2022NeurIPS89 citations

Proposes Conditional Mixture-of-Experts routing strategies for generalist models to resolve cross-task and cross-modality parameter interference, achieving state-of-the-art multi-task and zero-shot performance with minimal data and computational cost.

Listen

Artificial intelligence research has increasingly shifted toward unified generalist models that can execute multiple tasks across text, image, and video modalities using a single shared set of parameters. However, when a single model handles diverse tasks simultaneously, it often experiences performance degradation compared to specialized, single-task systems. This issue, known as task interference, arises because the optimization gradients required for distinct objectives frequently point in conflicting directions, leading to suboptimal compromises within shared model parameters.

The main objective of the article is to identify and quantify the mechanics of task interference in generalist models and to demonstrate that a conditional routing framework can mitigate cross-task conflicts without sacrificing the model's ability to adapt or generalize to unseen downstream tasks.

To achieve this, the authors analyzed gradient behavior across various layers of a generalist model and evaluated several routing configurations within a Mixture-of-Experts (MoE) architecture. Unlike conventional approaches that introduce task-specific components or route computations solely based on token patterns, the authors designed Conditional Mixture-of-Experts (Conditional MoEs) and tested them across five condition levels: token, context, modality, task, and predefined token attributes. The experimental framework integrated these techniques into the Uni-Perceiver model family (across Tiny, Base, and Large configurations) using multi-modal datasets—including ImageNet, Books&Wiki, MSCOCO, Kinetics-400, Flickr30k, and GLUE—to evaluate both standard and zero-shot novel task performance.

The evaluation yielded several critical findings. First, gradient conflict is significantly more pronounced in deeper model layers than in shallow layers, confirming that deep representations suffer the most from shared-parameter multi-task interference. Second, routing experts using an 8-bit attribute embedding—which encodes input/target modalities, causality, and token source—achieved superior performance over token-, context-, and modality-level variants while avoiding the generalization bottlenecks of task-specific IDs. Third, when incorporating attribute-conditioned experts, the generalist model matched or outperformed task-specialized baselines across vision, language, and cross-modal tasks. Fourth, when tuned using lightweight prompts on only 1% of downstream data, the model achieved performance competitive with state-of-the-art benchmarks while consuming less than 5% of the training data and under 10% of the training cost. Finally, the model retained strong generalization, establishing robust zero-shot and few-shot results on unseen tasks such as video captioning and video-text retrieval.

These findings indicate that sparse parameter activation conditioned on task and modality attributes substantially improves multi-task learning efficiency. By reducing data dependency and computing requirements, organizations can train and deploy highly capable foundation models at a fraction of the cost, timeline, and memory overhead typical of massive dense systems. Furthermore, using reparameterization techniques for data-independent attribute routing eliminates the excessive memory and inter-device communication latency typically associated with standard Mixture-of-Experts architectures during deployment.

Based on these results, engineering teams and organizations deploying multi-modal models should consider adopting attribute-based conditional routing as an architectural standard to resolve task conflict. Decision-makers should leverage lightweight prompt tuning rather than complete model fine-tuning on large downstream datasets to minimize compute expenditure. Before broad enterprise deployment on giant architectures, teams should run targeted pilot programs to evaluate whether the scaling benefits observed in hundred-million-parameter models hold at multi-billion-parameter scales.

The study's primary limitation is that it evaluates models scaled up to several hundred million parameters, leaving the exact behavior of Conditional MoEs on billion-scale foundation models an open empirical question. Additionally, like all models pre-trained on large-scale web datasets, the system carries potential societal concerns regarding energy consumption during training and latent dataset biases. Nonetheless, for the evaluated scales and tasks, the findings provide high confidence that conditional sparse routing effectively resolves multi-task degradation.

Cover for Uni-Perceiver-MoE: Learning Sparse Generalist Models with Conditional MoEs

Abstract

To build an artificial neural network like the biological intelligence system, recent works have unified numerous tasks into a generalist model, which can process various tasks with shared parameters and do not have any task-specific modules. While generalist models achieve promising results on various benchmarks, they have performance degradation on some tasks compared with task-specialized models. In this work, we find that interference among different tasks and modalities is the main factor to this phenomenon. To mitigate such interference, we introduce the Conditional Mixture-of-Experts (Conditional MoEs) to generalist models. Routing strategies under different levels of conditions are proposed to take both the training/inference cost and generalization ability into account. By incorporating the proposed Conditional MoEs, the recently proposed generalist model Uni-Perceiver can effectively mitigate the interference across tasks and modalities, and achieves state-of-the-art results on a series of downstream tasks via prompt tuning on 1% of downstream data. Moreover, the introduction of Conditional MoEs still holds the generalization ability of generalist models to conduct zero-shot inference on new tasks, e.g., video-text retrieval and video caption. Code and pre-trained generalist models are publicly released at https://github.com/fundamentalvision/Uni-Perceiver.

Table of Contents

  • 1 Introduction
  • 2 Related Works
  • 3 Methodology
  • 3.1 Task Interference
  • 3.2 Conditional Mixture-of-Experts (Conditional MoEs)
  • 3.3 Comparison of Conditional-MoE Variants
  • 4 Experiments
  • 4.1 Datasets
  • 4.2 Implementation Details
  • 4.3 Ablation Studies
  • 4.4 Evaluation on Pre-training tasks
  • 4.5 Generalization to Novel Tasks
  • 5 Conclusion
  • References
  • Checklist

Knowls

  1. Knowl 1 — Metric for Quantifying Inter-Task Gradient Interference

    equation

    To evaluate how training on task jj affects the optimization of task ii when sharing parameters θ\theta, the expected change in the loss function LiL_i of task ii induced by a step in the normalized gradient direction of task jj's loss LjL_j is approximated via a first-order Taylor expansion:

    ΔjLi(xi)≐Exj[Li(xi;θ)−Li(xi;θ−λ∇θLj(xj)∥∇θLj(xj)∥)]≈λExj[(∇θLj(xj)∥∇θLj(xj)∥)T∇θLi(xi)]\Delta_j L_i(x_i) \doteq \mathbb{E}_{x_j} \left[ L_i(x_i; \theta) - L_i\left(x_i; \theta - \lambda \frac{\nabla_\theta L_j(x_j)}{\|\nabla_\theta L_j(x_j)\|}\right) \right] \approx \lambda \mathbb{E}_{x_j} \left[ \left(\frac{\nabla_\theta L_j(x_j)}{\|\nabla_\theta L_j(x_j)\|}\right)^T \nabla_\theta L_i(x_i) \right]

    where xix_i and xjx_j are input batches sampled from task ii and task jj respectively, and λ>0\lambda > 0 is the learning rate step size. The normalized inter-task interference metric Ii,jI_{i,j} of task jj on task ii is defined as:

    Ii,j=Exi[ΔjLi(xi)ΔiLi(xi)]I_{i,j} = \mathbb{E}_{x_i} \left[ \frac{\Delta_j L_i(x_i)}{\Delta_i L_i(x_i)} \right]

    A value Ii,j>0I_{i,j} > 0 indicates that updating on task jj positively cooperates with reducing the loss of task ii, while Ii,j<0I_{i,j} < 0 indicates negative interference (conflicting update directions).

  2. Knowl 2 — Conditional Mixture-of-Experts Formulation for Generalist Models

    model/method

    Conditional Mixture-of-Experts (Conditional MoEs) replace standard linear projection layers in Transformer self-attention and feed-forward network (FFN) blocks with sparse expert routing conditioned on contextual, modal, task, or attribute signals.

    Given an input token xi∈Rdx_i \in \mathbb{R}^d from a sequence X={xi}i=1LX = \{x_i\}_{i=1}^L, a gating decision vector G∈REG \in \mathbb{R}^E over EE candidate expert linear projections is computed as:

    G=topk(softmax(Wg⋅R(xi)+ϵ))G = \text{top}_k\left(\text{softmax}(W_g \cdot R(x_i) + \epsilon)\right)

    where R(xi)R(x_i) denotes a conditioning routing representation of token xix_i, Wg∈RE×dRW_g \in \mathbb{R}^{E \times d_R} is a learnable gating projection matrix, ϵ\epsilon is standard normal exploration noise added to the gate logits, and the topk(⋅)\text{top}_k(\cdot) operator retains only the kk largest values (where k≪Ek \ll E) and sets all other E−kE - k entries to zero.

    The final output vector yi∈Rdouty_i \in \mathbb{R}^{d_{\text{out}}} is computed as the gate-weighted combination of the active expert linear transformations:

    yi=∑e=1EGe⋅We⋅xiy_i = \sum_{e=1}^E G_e \cdot W_e \cdot x_i

    where WeW_e denotes the weight matrix of the ee-th expert, and experts with Ge=0G_e = 0 are omitted from computation.

  3. Knowl 3 — 8-Dimensional Binary Token Attribute Scheme for Generalist Routing

    definition

    Attribute routing in Conditional MoEs assigns each input or target token xix_i an 8-dimensional binary vector attr(xi)∈{0,1}8\text{attr}(x_i) \in \{0, 1\}^8 describing the modality, causality, and role of the token without relying on explicit, task-specific IDs:

    • Index 0: Visual modality exists in the inputs of the current task (1=Yes1 = \text{Yes}, 0=No0 = \text{No}).
    • Index 1: Text modality exists in the inputs of the current task (1=Yes1 = \text{Yes}, 0=No0 = \text{No}).
    • Index 2: Visual modality exists in the targets of the current task (1=Yes1 = \text{Yes}, 0=No0 = \text{No}).
    • Index 3: Text modality exists in the targets of the current task (1=Yes1 = \text{Yes}, 0=No0 = \text{No}).
    • Index 4: The modality of the current token is visual (1=Yes1 = \text{Yes}, 0=No0 = \text{No}).
    • Index 5: The modality of the current token is text (1=Yes1 = \text{Yes}, 0=No0 = \text{No}).
    • Index 6: The attention mask applied to the current token is causal (1=Yes1 = \text{Yes}, 0=No0 = \text{No}).
    • Index 7: The current token comes from the task inputs (11) rather than the targets (00).

    For example, an input token for an image classification task (visual input to textual label targets) is assigned the attribute vector [1,0,0,1,1,0,0,1][1, 0, 0, 1, 1, 0, 0, 1].

    The routing conditioning vector is computed from the attribute vector by linear projection followed by Layer Normalization:

    Rattr(xi)=LayerNorm(Wattr⋅attr(xi))R_{\text{attr}}(x_i) = \text{LayerNorm}(W_{\text{attr}} \cdot \text{attr}(x_i))

    where WattrW_{\text{attr}} is a trainable transformation matrix.

  4. Knowl 4 — Conditioning Routing Strategies in Conditional MoEs

    model/method

    Conditional MoEs define alternative routing representations R(xi)R(x_i) for a token xix_i depending on the level of conditioning:

    1. Token-Level Routing (Data-Dependent): Uses the token's own feature vector directly: Rtoken(xi)=xiR_{\text{token}}(x_i) = x_i

    2. Context-Level Routing (Data-Dependent): Combines local token representation with global sequence context pooled across all LL tokens in XX: Rcontext(xi)=concat(xi,attnpool(X))R_{\text{context}}(x_i) = \text{concat}(x_i, \text{attnpool}(X))

    3. Modality-Level Routing (Data-Independent): Maps the modality index of the token through a modality embedding lookup: Rmodal(xi)=embed(idmodal(xi))R_{\text{modal}}(x_i) = \text{embed}(\text{id}_{\text{modal}}(x_i))

    4. Task-Level Routing (Data-Independent): Maps the unique task index through a task embedding lookup, routing all tokens in a given task to the same expert subset: Rtask(xi)=embed(idtask(xi))R_{\text{task}}(x_i) = \text{embed}(\text{id}_{\text{task}}(x_i))

    5. Attribute-Level Routing (Data-Independent): Projects an 8-bit binary descriptor attr(xi)\text{attr}(x_i) encoding input/target modalities, token modality, causality, and input/target status: Rattr(xi)=LayerNorm(Wattr⋅attr(xi))R_{\text{attr}}(x_i) = \text{LayerNorm}(W_{\text{attr}} \cdot \text{attr}(x_i))

  5. Knowl 5 — Reparameterization and Efficiency of Data-Independent Conditional MoEs

    model/method

    Conditional MoE variants differ fundamentally in memory footprint and inference cost based on whether gating decisions depend on dynamically varying token representations (data-dependent) or static descriptors (data-independent):

    • Data-dependent variants (token-level and context-level routing) evaluate routing dynamically for each token, requiring all expert weights to remain loaded in device memory and necessitating expert parallelism across distributed GPUs during both training and inference.
    • Data-independent variants (modality-level, task-level, and attribute-level routing) produce identical gating weights GG for all tokens that share the same condition. Consequently, the active expert projection weights can be pre-combined via reparameterization: Wmerged=∑e=1EGe⋅WeW_{\text{merged}} = \sum_{e=1}^E G_e \cdot W_e This allows the MoE layer to collapse into a standard single dense linear projection (yi=Wmerged⋅xiy_i = W_{\text{merged}} \cdot x_i), reducing the computational and memory cost during inference to that of a dense non-MoE network.
  6. Knowl 6 — Layer-Wise Divergence in Multi-Task Gradient Interference

    empirical result

    Empirical evaluation of the inter-task gradient interference metric Ii,jI_{i,j} across layers of a multi-task pre-trained generalist model (Uni-Perceiver-Ti) shows that task synergy occurs primarily in shallow layers, while destructive task interference dominates deep layers.

    4-th FFN Block (Ii,jI_{i,j}) 12-th FFN Block (Ii,jI_{i,j})
    Task ii ask jj ImgCLS MLM Caption ImgCLS MLM Caption
    ImgCLS (Img) 1.00 -0.57 1.29 1.00 -2.91 -2.45
    MLM (Text) 0.07 1.00 0.68 -1.65 1.00 -1.05
    Caption (Img-Text) 0.01 0.01 1.00 -0.11 0.19 1.00

    At the 4th FFN block, Image Captioning gradients positively impact Image Classification (I=1.29I = 1.29) and Masked Language Modeling (I=0.68I = 0.68). By the 12th FFN block, cross-task gradients become sharply antagonistic (e.g., MLM on Image Classification drops to I=−2.91I = -2.91, and Image Captioning on Image Classification drops to I=−2.45I = -2.45).

  7. Knowl 7 — Ablation of Routing Strategies on Multi-Task Performance and Efficiency

    empirical result

    Ablation of routing conditioning methods on Uni-Perceiver-Ti across ImageNet-1k classification, COCO Captioning, and Books&Wiki Masked Language Modeling (MLM) shows that attribute routing achieves the best trade-off between task performance, training overhead, and inference speed.

    Routing Strategy Train Infer. ImageNet-1k (%) COCO Caption MLM
    Time Time acctrain\text{acc}_{\text{train}} accval\text{acc}_{\text{val}} acctrain\text{acc}_{\text{train}} B@4val\text{B@4}_{\text{val}} acctrain\text{acc}_{\text{train}} pplval\text{ppl}_{\text{val}}
    Fully Shared (Uni-Perceiver-Ti) 1.0×1.0\times 1.0×1.0\times 47.3 68.3 49.2 18.2 54.5 5.86
    Task-Specific Parameters 1.1×1.1\times 1.0×1.0\times 53.3 73.5 52.6 20.4 60.5 4.48
    Conditional MoEs (Token-level) 1.8×1.8\times 2.2×2.2\times 53.1 72.7 52.9 20.9 58.3 4.96
    Conditional MoEs (Context-level) 2.2×2.2\times 2.6×2.6\times 52.5 73.1 52.8 21.5 58.6 4.86
    Conditional MoEs (Modality-level) 1.4×1.4\times 1.0×1.0\times 51.7 72.6 52.1 21.8 57.5 5.06
    Conditional MoEs (Task-level) 1.4×1.4\times 1.0×1.0\times 52.9 73.2 52.7 21.2 59.9 4.56
    Conditional MoEs (Attribute-level) 1.4×1.4\times 1.0×1.0\times 52.8 73.3 53.1 23.0 60.0 4.56

    Attribute-level routing recovers the performance loss of parameter sharing without requiring task-specific parameter allocation, outperforming vanilla Uni-Perceiver-Ti by +5.0% validation accuracy on ImageNet-1k, +4.8 BLEU@4 on COCO Caption, and reducing validation perplexity from 5.86 to 4.56 on Books&Wiki.

  8. Knowl 8 — Downstream Performance of Uni-Perceiver-MoE on Pre-Trained Modalities

    empirical result

    Integrating Attribute Conditional MoEs into Uni-Perceiver Base (B) and Large (L) backbones brings consistent gains across visual classification and multimodal retrieval benchmarks, evaluated without tuning (WT), with 1% downstream prompt tuning (PT1%), and with 100% fine-tuning (FT100%):

    • ImageNet-1k Top-1 Accuracy (%):
      • Uni-Perceiver-B: WT = 79.2, PT1% = 80.9, FT100% = 84.0
      • Uni-Perceiver-B + Cond-MoE: WT = 80.3, PT1% = 82.0, FT100% = 84.5
      • Uni-Perceiver-L: WT = 82.7, PT1% = 84.2, FT100% = 86.2
      • Uni-Perceiver-L + Cond-MoE: WT = 83.4, PT1% = 84.9, FT100% = 86.4
    • Kinetics-400 Top-1 Accuracy (%):
      • Uni-Perceiver-B: WT = 74.5, PT1% = 74.8, FT100% = 77.7
      • Uni-Perceiver-B + Cond-MoE: WT = 76.8, PT1% = 77.2, FT100% = 79.3
      • Uni-Perceiver-L: WT = 79.5, PT1% = 80.0, FT100% = 81.9
      • Uni-Perceiver-L + Cond-MoE: WT = 82.1, PT1% = 83.0, FT100% = 84.2
    • Flickr30k Retrieval (Recall@1 %):
      • Uni-Perceiver-L (Image →\to Text): WT = 83.7, PT1% = 92.1, FT100% = 94.7
      • Uni-Perceiver-L + Cond-MoE (Image →\to Text): WT = 83.6, PT1% = 92.4, FT100% = 94.1
      • Uni-Perceiver-L (Text →\to Image): WT = 74.2, PT1% = 80.0, FT100% = 82.1
      • Uni-Perceiver-L + Cond-MoE (Text →\to Image): WT = 75.9, PT1% = 80.6, FT100% = 83.7
    • MSCOCO Image Captioning (BLEU@4):
      • Uni-Perceiver-B: WT = 32.0, PT1% = 35.5, FT100% = 36.4
      • Uni-Perceiver-B + Cond-MoE: WT = 33.2, PT1% = 36.8, FT100% = 37.3
      • Uni-Perceiver-L: WT = 35.3, PT1% = 38.6, FT100% = 39.2
      • Uni-Perceiver-L + Cond-MoE: WT = 35.5, PT1% = 39.3, FT100% = 40.5
  9. Knowl 9 — Zero-Shot and Few-Shot Generalization to Unseen Tasks (GLUE and MSVD)

    empirical result

    Conditional MoEs preserve and enhance the zero-shot and few-shot transferability of generalist models when evaluated on novel downstream tasks not encountered during multi-task pre-training:

    1. Natural Language Understanding (GLUE Benchmark, Fine-Tuned):

      • Uni-Perceiver-B vs Uni-Perceiver-B + Cond-MoE: MNLI improves from 79.7% to 81.5%, QNLI from 87.3% to 88.2%, QQP from 86.7% to 87.8%, RTE from 71.1% to 75.8%, SST-2 from 89.3% to 90.9%, MRPC from 86.0% to 87.1%, and CoLA from 43.1% to 52.2% Mcc.
      • Uni-Perceiver-L + Cond-MoE achieves 85.7% (MNLI), 91.9% (QNLI), 89.5% (QQP), 78.4% (RTE), 93.4% (SST-2), 91.2% (MRPC), and 57.4% (CoLA).
    2. Video-Text Retrieval and Video Captioning (MSVD Benchmark):

      • Uni-Perceiver-B + Cond-MoE achieves zero-shot (WT) Video →\to Text R@1 of 52.8% (vs 50.3% vanilla), Text →\to Video R@1 of 40.0% (vs 38.7% vanilla), and Video Captioning BLEU@4 of 23.4% (vs 22.6% vanilla).
      • With 1% prompt tuning (PT1%), Uni-Perceiver-L + Cond-MoE reaches 66.4% Video →\to Text R@1, 50.3% Text →\to Video R@1, and 67.6 BLEU@4, performing on par with models trained with specialized architectures.
  10. Knowl 10 — Scale Limitation of Conditional MoEs for Generalist Models

    limitation

    The effectiveness of Conditional MoEs in mitigating multi-task interference is verified empirically on generalist model scales up to hundreds of millions of parameters (Uni-Perceiver-Ti, Base, and Large, with deployed parameter counts between 86M and 505M). Whether task-interference exhibits identical characteristics or whether Conditional MoEs provide comparable benefits in generalist models scaled to billions of parameters remains unverified.

Coverage note — None was omitted; all core methodology, routing variants, theoretical formulation of interference, empirical tables, downstream benchmark results, and stated limitations are fully represented.

References

  1. 1.A. Abbas and Y. Andreopoulos. Biased mixtures of experts: Enabling computer vision inference under data transfer limitations. TIP, 2020.
  2. 2.H. Akbari, L. Yuan, R. Qian, W.-H. Chuang, S.-F. Chang, Y. Cui, and B. Gong. Vatt: Transformers for multimodal self-supervised learning from raw video, audio and text. NIPS, 2021.
  3. 3.J.-B. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y. Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds, et al. Flamingo: a visual language model for few-shot learning. arXiv preprint arXiv:2204.14198, 2022.
  4. 4.N. Arivazhagan, A. Bapna, O. Firat, D. Lepikhin, M. Johnson, M. Krikun, M. X. Chen, Y. Cao, G. Foster, C. Cherry, et al. Massively multilingual neural machine translation in the wild: Findings and challenges. arXiv preprint arXiv:1907.05019, 2019.
  5. 5.A. Arnab, M. Dehghani, G. Heigold, C. Sun, M. Lu\v{c}i'c, and C. Schmid. Vivit: A video vision transformer. arXiv preprint arXiv:2103.15691, 2021.
  6. 6.J. L. Ba, J. R. Kiros, and G. E. Hinton. Layer normalization. arXiv preprint arXiv:1607.06450, 2016.
  7. 7.G. Bertasius, H. Wang, and L. Torresani. Is space-time attention all you need for video understanding? arXiv preprint arXiv:2102.05095, 2021.
  8. 8.R. Caruana. Multitask learning. Machine learning, 1997.
  9. 9.S. Changpinyo, P. Sharma, N. Ding, and R. Soricut. Conceptual 12M: Pushing web-scale image-text pre-training to recognize long-tail visual concepts. In CVPR, 2021.
  10. 10.D. S. Chaplot, L. Lee, R. Salakhutdinov, D. Parikh, and D. Batra. Embodied multimodal multitask learning. arXiv preprint arXiv:1902.01385, 2019.
  11. 11.D. Chen and W. B. Dolan. Collecting highly parallel data for paraphrase evaluation. In ACL, 2011.
  12. 12.X. Chen, H. Fang, T.-Y. Lin, R. Vedantam, S. Gupta, P. Doll'ar, and C. L. Zitnick. Microsoft coco captions: Data collection and evaluation server. arXiv preprint arXiv:1504.00325, 2015.
  13. 13.Y.-C. Chen, L. Li, L. Yu, A. El Kholy, F. Ahmed, Z. Gan, Y. Cheng, and J. Liu. Uniter: Universal image-text representation learning. 2020.
  14. 14.Z. Chen, V. Badrinarayanan, C.-Y. Lee, and A. Rabinovich. Gradnorm: Gradient normalization for adaptive loss balancing in deep multitask networks. In ICML, 2018.
  15. 15.B. Cheng, I. Misra, A. G. Schwing, A. Kirillov, and R. Girdhar. Masked-attention mask transformer for universal image segmentation. 2022.
  16. 16.K. Clark, M.-T. Luong, U. Khandelwal, C. D. Manning, and Q. V. Le. Bam! born-again multi-task networks for natural language understanding. arXiv preprint arXiv:1907.04829, 2019.
  17. 17.M. Crawshaw. Multi-task learning with deep neural networks: A survey. arXiv preprint arXiv:2009.09796, 2020.
  18. 18.J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei. Imagenet: A large-scale hierarchical image database. In CVPR, 2009.
  19. 19.J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018.
  20. 20.A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929, 2020.
  21. 21.N. Du, Y. Huang, A. M. Dai, S. Tong, D. Lepikhin, Y. Xu, M. Krikun, Y. Zhou, A. W. Yu, O. Firat, et al. Glam: Efficient scaling of language models with mixture-of-experts. arXiv preprint arXiv:2112.06905, 2021.
  22. 22.H. Fang, P. Xiong, L. Xu, and Y. Chen. Clip2video: Mastering video-text retrieval via image clip. arXiv preprint arXiv:2106.11097, 2021.
  23. 23.W. Fedus, B. Zoph, and N. Shazeer. Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity. arXiv preprint arXiv:2101.03961, 2021.
  24. 24.M. Guo, A. Haque, D.-A. Huang, S. Yeung, and L. Fei-Fei. Dynamic task prioritization for multitask learning. In ECCV, 2018.
  25. 25.K. Hashimoto, C. Xiong, Y. Tsuruoka, and R. Socher. A joint many-task model: Growing a neural network for multiple nlp tasks. arXiv preprint arXiv:1611.01587, 2016.
  26. 26.K. He, X. Zhang, S. Ren, and J. Sun. Deep residual learning for image recognition. In CVPR, 2016.
  27. 27.K. He, G. Gkioxari, P. Doll'ar, and R. Girshick. Mask r-cnn. In ICCV, 2017.
  28. 28.C. Hokamp, J. Glover, and D. Gholipour. Evaluating the supervised and zero-shot performance of multi-lingual translation models. arXiv preprint arXiv:1906.09675, 2019.
  29. 29.R. Hu and A. Singh. Unit: Multimodal multitask learning with a unified transformer. arXiv preprint arXiv:2102.10772, 2021.
  30. 30.T. Iki and A. Aizawa. Effect of visual extensions on natural language understanding in vision-and-language models. arXiv preprint arXiv:2104.08066, 2021.
  31. 31.A. Jaegle, S. Borgeaud, J.-B. Alayrac, C. Doersch, C. Ionescu, D. Ding, S. Koppula, A. Brock, E. Shelhamer, O. H'enaff, M. M. Botvinick, A. Zisserman, O. Vinyals, and J. Carreira. Perceiver io: A general architecture for structured inputs & outputs, 2021.
  32. 32.A. Jaegle, F. Gimeno, A. Brock, O. Vinyals, A. Zisserman, and J. Carreira. Perceiver: General perception with iterative attention. In ICML, 2021.
  33. 33.C. Jia, Y. Yang, Y. Xia, Y.-T. Chen, Z. Parekh, H. Pham, Q. Le, Y.-H. Sung, Z. Li, and T. Duerig. Scaling up visual and vision-language representation learning with noisy text supervision. In International Conference on Machine Learning, pages 4904--4916. PMLR, 2021.
  34. 34.H. Jiang, P. He, W. Chen, X. Liu, J. Gao, and T. Zhao. Smart: Robust and efficient fine-tuning for pre-trained natural language models through principled regularized optimization. arXiv preprint arXiv:1911.03437, 2019.
  35. 35.S. Kalkowski, C. Schulze, A. Dengel, and D. Borth. Real-time analysis and visualization of the yfcc100m dataset. In Proceedings of the 2015 workshop on community-organized multimodal mining: opportunities for novel solutions, pages 25--30, 2015.
  36. 36.M. Kanakis, D. Bruggemann, S. Saha, S. Georgoulis, A. Obukhov, and L. V. Gool. Reparameterizing convolutions for incremental multi-task learning without task interference. 2020.
  37. 37.W. Kay, J. Carreira, K. Simonyan, B. Zhang, C. Hillier, S. Vijayanarasimhan, F. Viola, T. Green, T. Back, P. Natsev, et al. The kinetics human action video dataset. arXiv preprint arXiv:1705.06950, 2017.
  38. 38.A. Kendall, Y. Gal, and R. Cipolla. Multi-task learning using uncertainty to weigh losses for scene geometry and semantics. In CVPR, 2018.
  39. 39.W. Kim, B. Son, and I. Kim. Vilt: Vision-and-language transformer without convolution or region supervision. arXiv preprint arXiv:2102.03334, 2021.
  40. 40.Y. J. Kim, A. A. Awan, A. Muzio, A. F. C. Salinas, L. Lu, A. Hendy, S. Rajbhandari, Y. He, and H. H. Awadalla. Scalable and efficient moe training for multitask multilingual models. arXiv preprint arXiv:2109.10465, 2021.
  41. 41.R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L.-J. Li, D. A. Shamma, et al. Visual genome: Connecting language and vision using crowdsourced dense image annotations. IJCV, 123(1):32--73, 2017.
  42. 42.S. Kudugunta, Y. Huang, A. Bapna, M. Krikun, D. Lepikhin, M.-T. Luong, and O. Firat. Beyond distillation: Task-level mixture-of-experts for efficient inference. arXiv preprint arXiv:2110.03742, 2021.
  43. 43.D. Lepikhin, H. Lee, Y. Xu, D. Chen, O. Firat, Y. Huang, M. Krikun, N. Shazeer, and Z. Chen. Gshard: Scaling giant models with conditional computation and automatic sharding. arXiv preprint arXiv:2006.16668, 2020.
  44. 44.J. Li, D. Li, C. Xiong, and S. Hoi. Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation. arXiv preprint arXiv:2201.12086, 2022.
  45. 45.L. H. Li, M. Yatskar, D. Yin, C.-J. Hsieh, and K.-W. Chang. Visualbert: A simple and performant baseline for vision and language. arXiv preprint arXiv:1908.03557, 2019.
  46. 46.X. Li, X. Yin, C. Li, P. Zhang, X. Hu, L. Zhang, L. Wang, H. Hu, L. Dong, F. Wei, et al. Oscar: Object-semantics aligned pre-training for vision-language tasks. In ECCV, 2020.
  47. 47.Z. Li, F. Liu, W. Yang, S. Peng, and J. Zhou. A survey of convolutional neural networks: Analysis, applications, and prospects. IEEE Transactions on Neural Networks and Learning Systems, pages 1--21, 2021. doi: 10.1109/TNNLS.2021.3084827.
  48. 48.Z. Li, W. Wang, E. Xie, Z. Yu, A. Anandkumar, J. Alvarez, T. Lu, and P. Luo. Panoptic segformer: Delving deeper into panoptic segmentation with transformers. 2022.
  49. 49.Z. Lin, L. Wu, M. Wang, and L. Li. Learning language specific sub-network for multilingual machine translation. arXiv preprint arXiv:2105.09259, 2021.
  50. 50.P. Liu, X. Qiu, and X. Huang. Adversarial multi-task learning for text classification. arXiv preprint arXiv:1704.05742, 2017.
  51. 51.X. Liu, Y. Zheng, Z. Du, M. Ding, Y. Qian, Z. Yang, and J. Tang. Gpt understands, too. arXiv preprint arXiv:2103.10385, 2021.
  52. 52.Y. Liu, M. Ott, N. Goyal, J. Du, M. Joshi, D. Chen, O. Levy, M. Lewis, L. Zettlemoyer, and V. Stoyanov. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692, 2019.
  53. 53.Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo. Swin transformer: Hierarchical vision transformer using shifted windows. ICCV, 2021.
  54. 54.J. Lu, D. Batra, D. Parikh, and S. Lee. Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks. arXiv preprint arXiv:1908.02265, 2019.
  55. 55.J. Lu, V. Goswami, M. Rohrbach, D. Parikh, and S. Lee. 12-in-1: Multi-task vision and language representation learning. In CVPR, 2020.
  56. 56.S. Min, W. Kong, R.-C. Tu, D. Gong, C. Cai, W. Zhao, C. Liu, S. Zheng, H. Wang, Z. Li, et al. Hunyuan_tvr for text-video retrivial. arXiv preprint arXiv:2204.03382, 2022.
  57. 57.M. Monfort, A. Andonian, B. Zhou, K. Ramakrishnan, S. A. Bargal, T. Yan, L. Brown, Q. Fan, D. Gutfreund, C. Vondrick, et al. Moments in time dataset: one million videos for event understanding. TPAMI, 2019.
  58. 58.V. Ordonez, G. Kulkarni, and T. Berg. Im2text: Describing images using 1 million captioned photographs. NeurIPS, 2011.
  59. 59.B. A. Plummer, L. Wang, C. M. Cervantes, J. C. Caicedo, J. Hockenmaier, and S. Lazebnik. Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models. In ICCV, 2015.
  60. 60.D. Qi, L. Su, J. Song, E. Cui, T. Bharti, and A. Sacheti. Imagebert: Cross-modal pre-training with large-scale weak-supervised image-text data. arXiv preprint arXiv:2001.07966, 2020.
  61. 61.A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al. Learning transferable visual models from natural language supervision. arXiv preprint arXiv:2103.00020, 2021.
  62. 62.S. Reed, K. Zolna, E. Parisotto, S. G. Colmenarejo, A. Novikov, G. Barth-Maron, M. Gimenez, Y. Sulsky, J. Kay, J. T. Springenberg, T. Eccles, J. Bruce, A. Razavi, A. Edwards, N. Heess, Y. Chen, R. Hadsell, O. Vinyals, M. Bordbar, and N. de Freitas. A generalist agent, 2022.
  63. 63.S. Ren, K. He, R. Girshick, and J. Sun. Faster r-cnn: Towards real-time object detection with region proposal networks. NeurIPS, 2015.
  64. 64.C. Riquelme, J. Puigcerver, B. Mustafa, M. Neumann, R. Jenatton, A. Susano Pinto, D. Keysers, and N. Houlsby. Scaling vision with sparse mixture of experts. NIPS, 2021.
  65. 65.J. Shao, S. Chen, Y. Li, K. Wang, Z. Yin, Y. He, J. Teng, Q. Sun, M. Gao, J. Liu, et al. Intern: A new learning paradigm towards general vision. arXiv preprint arXiv:2111.08687, 2021.
  66. 66.P. Sharma, N. Ding, S. Goodman, and R. Soricut. Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning. In ACL, 2018.
  67. 67.N. Shazeer, A. Mirhoseini, K. Maziarz, A. Davis, Q. Le, G. Hinton, and J. Dean. Outrageously large neural networks: The sparsely-gated mixture-of-experts layer. arXiv preprint arXiv:1701.06538, 2017.
  68. 68.N. Shazeer, Y. Cheng, N. Parmar, D. Tran, A. Vaswani, P. Koanantakool, P. Hawkins, H. Lee, M. Hong, C. Young, et al. Mesh-tensorflow: Deep learning for supercomputers. NIPS, 31, 2018.
  69. 69.S. Shen, L. H. Li, H. Tan, M. Bansal, A. Rohrbach, K.-W. Chang, Z. Yao, and K. Keutzer. How much can clip benefit vision-and-language tasks? arXiv preprint arXiv:2107.06383, 2021.
  70. 70.K. Simonyan and A. Zisserman. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556, 2014.
  71. 71.A. Singh, R. Hu, V. Goswami, G. Couairon, W. Galuba, M. Rohrbach, and D. Kiela. Flava: A foundational language and vision alignment model. arXiv preprint arXiv:2112.04482, 2021.
  72. 72.T. Standley, A. Zamir, D. Chen, L. Guibas, J. Malik, and S. Savarese. Which tasks should be learned together in multi-task learning? In ICML, 2020.
  73. 73.A. Steiner, A. Kolesnikov, X. Zhai, R. Wightman, J. Uszkoreit, and L. Beyer. How to train your vit? data, augmentation, and regularization in vision transformers. arXiv preprint arXiv:2106.10270, 2021.
  74. 74.G. Strezoski, N. v. Noord, and M. Worring. Many task learning with task routing. In ICCV, 2019.
  75. 75.H. Tan and M. Bansal. Lxmert: Learning cross-modality encoder representations from transformers. arXiv preprint arXiv:1908.07490, 2019.
  76. 76.H. Touvron, M. Cord, M. Douze, F. Massa, A. Sablayrolles, and H. J'egou. Training data-efficient image transformers & distillation through attention. In ICML, 2021.
  77. 77.A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, \L. Kaiser, and I. Polosukhin. Attention is all you need. In NeurIPS, 2017.
  78. 78.A. Wang, A. Singh, J. Michael, F. Hill, O. Levy, and S. R. Bowman. Glue: A multi-task benchmark and analysis platform for natural language understanding. arXiv preprint arXiv:1804.07461, 2018.
  79. 79.P. Wang, A. Yang, R. Men, J. Lin, S. Bai, Z. Li, J. Ma, C. Zhou, J. Zhou, and H. Yang. Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework. arXiv preprint arXiv:2202.03052, 2022.
  80. 80.W. Wang, E. Xie, X. Li, D.-P. Fan, K. Song, D. Liang, T. Lu, P. Luo, and L. Shao. Pyramid vision transformer: A versatile backbone for dense prediction without convolutions. In ICCV, 2021.
  81. 81.X. Wang, Y. Tsvetkov, and G. Neubig. Balancing training for multilingual neural machine translation. arXiv preprint arXiv:2004.06748, 2020.
  82. 82.X. Wang, F. Yu, L. Dunlap, Y.-A. Ma, R. Wang, A. Mirhoseini, T. Darrell, and J. E. Gonzalez. Deep mixture of experts via shallow embedding. In Uncertainty in artificial intelligence. PMLR, 2020.
  83. 83.Z. Wang, Z. C. Lipton, and Y. Tsvetkov. On negative interference in multilingual models: Findings and a meta-learning treatment. arXiv preprint arXiv:2010.03017, 2020.
  84. 84.Z. Wang, J. Yu, A. W. Yu, Z. Dai, Y. Tsvetkov, and Y. Cao. Simvlm: Simple visual language model pretraining with weak supervision. arXiv preprint arXiv:2108.10904, 2021.
  85. 85.B. Yang, G. Bender, Q. V. Le, and J. Ngiam. Condconv: Conditionally parameterized convolutions for efficient inference. NIPS, 32, 2019.
  86. 86.Z. Yang, Z. Gan, J. Wang, X. Hu, F. Ahmed, Z. Liu, Y. Lu, and L. Wang. Crossing the format boundary of text and boxes: Towards unified vision-language modeling. arXiv preprint arXiv:2111.12085, 2021.
  87. 87.J. Yu, Z. Wang, V. Vasudevan, L. Yeung, M. Seyedhosseini, and Y. Wu. Coca: Contrastive captioners are image-text foundation models. arXiv preprint arXiv:2205.01917, 2022.
  88. 88.T. Yu, S. Kumar, A. Gupta, S. Levine, K. Hausman, and C. Finn. Gradient surgery for multi-task learning. NIPS, 33:5824--5836, 2020.
  89. 89.L. Yuan, D. Chen, Y.-L. Chen, N. Codella, X. Dai, J. Gao, H. Hu, X. Huang, B. Li, C. Li, et al. Florence: A new foundation model for computer vision. arXiv preprint arXiv:2111.11432, 2021.
  90. 90.B. Zhang, A. Bapna, R. Sennrich, and O. Firat. Share or not? learning to schedule language-specific capacity for multilingual translation. In ICLR, 2020.
  91. 91.Z. Zhang, Y. Shi, C. Yuan, B. Li, P. Wang, W. Hu, and Z.-J. Zha. Object relational graph with teacher-recommended learning for video captioning. In CVPR, pages 13278--13288, 2020.
  92. 92.L. Zhou, H. Palangi, L. Zhang, H. Hu, J. Corso, and J. Gao. Unified vision-language pre-training for image captioning and vqa. In AAAI, 2020.
  93. 93.X. Zhu, J. Zhu, H. Li, X. Wu, X. Wang, H. Li, X. Wang, and J. Dai. Uni-perceiver: Pre-training unified architecture for generic perception for zero-shot and few-shot tasks. arXiv preprint arXiv:2112.01522, 2021.
  94. 94.Y. Zhu, R. Kiros, R. Zemel, R. Salakhutdinov, R. Urtasun, A. Torralba, and S. Fidler. Aligning books and movies: Towards story-like visual explanations by watching movies and reading books. In ICCV, pages 19--27, 2015.

Citation

MLA
Zhu, J., et al. “Uni-Perceiver-MoE: Learning Sparse Generalist Models with Conditional MoEs”. Advances in Neural Information Processing Systems, vol. 35, 2022, pp. 2664–78, https://proceedings.neurips.cc/paper_files/paper/2022/file/11fc8c98b46d4cbdfe8157267228f7d7-Paper-Conference.pdf.
APA
Zhu, J., Zhu, X., Wang, W., Wang, X., Li, H., Wang, X., & Dai, J. (2022). Uni-Perceiver-MoE: Learning Sparse Generalist Models with Conditional MoEs. Advances in Neural Information Processing Systems, 35, 2664–2678. https://proceedings.neurips.cc/paper_files/paper/2022/file/11fc8c98b46d4cbdfe8157267228f7d7-Paper-Conference.pdf
Chicago
Zhu, J., X. Zhu, W. Wang, et al. 2022. “Uni-Perceiver-MoE: Learning Sparse Generalist Models with Conditional MoEs”. Advances in Neural Information Processing Systems 35: 2664–78. https://proceedings.neurips.cc/paper_files/paper/2022/file/11fc8c98b46d4cbdfe8157267228f7d7-Paper-Conference.pdf.
Harvard
Zhu, J. et al. (2022) “Uni-Perceiver-MoE: Learning Sparse Generalist Models with Conditional MoEs”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 2664–2678. Available at: https://proceedings.neurips.cc/paper_files/paper/2022/file/11fc8c98b46d4cbdfe8157267228f7d7-Paper-Conference.pdf.
Vancouver
1. Zhu J, Zhu X, Wang W, Wang X, Li H, Wang X, Dai J (2022) Uni-Perceiver-MoE: Learning Sparse Generalist Models with Conditional MoEs. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 2664–2678

BibTeX

@inproceedings{zhu2022uni,
  title = {Uni-Perceiver-MoE: Learning Sparse Generalist Models with Conditional MoEs},
  author = {Zhu, Jinguo and Zhu, Xizhou and Wang, Wenhai and Wang, Xiaohua and Li, Hongsheng and Wang, Xiaogang and Dai, Jifeng},
  year = {2022},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {35},
  pages = {2664-2678},
  url = {https://proceedings.neurips.cc/paper_files/paper/2022/file/11fc8c98b46d4cbdfe8157267228f7d7-Paper-Conference.pdf}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors