MoH: Multi-Head Attention as Mixture-of-Head Attention
Peng JinBo ZhuLi YuanShuicheng Yan
Proposes Mixture-of-Head attention, a parameter-neutral architecture that dynamically routes tokens to a subset of attention heads using weighted summation to cut computation while improving accuracy across vision models, diffusion models, and large language models.
Modern artificial intelligence models rely heavily on the Transformer architecture, where multi-head attention serves as the core computational mechanism. However, standard multi-head attention activates every attention head uniformly for every piece of data, despite growing evidence that many heads are redundant. This full activation introduces unnecessary computational overhead and raises inference costs when deploying large models at scale.
The article introduces Mixture-of-Head attention (MoH), a dynamic mechanism designed to reduce computational demands during inference while maintaining or exceeding baseline model accuracy, all without increasing total model parameters.
MoH treats individual attention heads as specialized experts within a dynamic routing framework. For each token, the system selects only a subset of the most relevant heads to activate, while a designated group of shared heads remains constantly active to retain common foundational knowledge. The architecture also incorporates a two-stage routing mechanism to weight head contributions and a load-balancing loss to keep heads evenly trained. The authors evaluated MoH across three mainstream AI domains: vision models for image classification (Vision Transformers), generative image diffusion models (Diffusion Transformers), and large language models (trained from scratch and continue-tuned on an existing 8-billion-parameter foundation model).
The evaluation produced several key findings. First, MoH models consistently matched or outperformed standard multi-head baselines while activating only 50% to 90% of total attention heads across vision, diffusion, and language domains. Second, existing foundation models can be converted into MoH models efficiently; continue-tuning a baseline 8-billion-parameter model with about 3% of its original pre-training budget produced a 2.4% average accuracy improvement across 14 benchmarks while using only 75% of heads. Third, dynamic routing reduced empirical inference latency, demonstrating greater time savings as sequence lengths expanded. Finally, ablation studies showed that always-active shared heads and two-stage routing are critical to preventing performance degradation.
These findings indicate that artificial intelligence architectures can achieve higher efficiency and specialized parameter use without parameter expansion or costly training from scratch. For organizations deploying generative and visual AI, adopting MoH offers a viable path to lower computational hardware costs, accelerate response latencies, and reduce energy expenses during inference. The successful conversion of pre-trained foundation models also provides an economical upgrade path that avoids multi-million-dollar re-training cycles.
Organizations evaluating this architecture should consider piloting MoH in high-throughput inference pipelines where latency and serving costs are critical. Implementers should maintain a relatively balanced proportion of shared heads (recommended at greater than 40% of activated heads) to ensure stability. Dense prediction tasks like image generation require conservative head reduction (e.g., activating around 90% of heads) compared to language or classification tasks (50% to 75%). Future work should explore extending head sparsity below 50%, scaling evaluations to larger parameter regimes beyond 8 billion, and testing performance on multimodal and audio architectures.
Confidence in these findings is high for classification and standard language benchmarks up to the 8-billion-parameter scale under controlled experimental conditions. However, decision-makers should note certain limitations: converting models on new data distributions showed slight performance drops in non-English and math tasks due to catastrophic forgetting during tuning. Real-world cost benefits will also depend on low-level software kernel support for sparse operations.
- Paper: Are Sixteen Heads Really Better than One?, Paul Michel et al. (2019). This study establishes that many attention heads can be removed at inference with little accuracy loss, motivating MoH’s selective activation of heads rather than uniform computation.
- Paper: Analyzing Multi-Head Self-Attention: Specialized Heads Do the Heavy Lifting, the Rest Can Be Pruned, Elena Voita et al. (2019). Its analysis of specialized and redundant heads provides the key background for understanding why MoH routes different tokens to different attention heads.
- Paper: Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer, Noam Shazeer et al. (2017). The sparsely gated Mixture-of-Experts layer introduces the conditional expert routing framework that MoH adapts from subnetworks to attention heads.
- Paper: Deja Vu: Contextual Sparsity for Efficient LLMs at Inference Time, Zichang Liu et al. (2023). DEJAVU demonstrates input-dependent sparsity in attention heads and predicts it for faster inference, a direct antecedent to MoH’s per-token head selection.
- Paper: AdaViT: Adaptive Vision Transformers for Efficient Image Recognition, Lingchen Meng et al. (2022). AdaViT’s adaptive selection of attention heads at inference provides a concrete vision-transformer precedent for MoH’s dynamic head activation.
No sufficiently relevant recommendations were found.
