MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark

S. SakshiUtkarsh TyagiSonal KumarAshish SethRamaneswaran SelvakumarOriol NietoRamani DuraiswamiSreyan GhoshDinesh Manocha

article2025ICLR217 citations

Presents MMAU, a 10,000-clip audio benchmark spanning speech, music, and environmental sounds across 27 complex reasoning skills, exposing severe capability gaps in state-of-the-art audio-language models that currently plateau near 53% accuracy.

Listen

Artificial intelligence systems increasingly excel at complex language and vision tasks, but their ability to process and reason about auditory information remains largely untested at an advanced level. Existing audio benchmarks focus primarily on basic perceptual capabilities, such as simple speech transcription or sound classification, which do not reflect the higher-order reasoning required for expert-level decision-making. Developing models that can effectively comprehend and interpret speech, environmental sounds, and music is essential for deploying reliable multimodal AI agents in real-world environments.

The article introduces and evaluates the Massive Multi-Task Audio Understanding and Reasoning Benchmark, known as MMAU. The primary objective is to evaluate the advanced audio perception, multi-step reasoning, and expert domain-knowledge retrieval capabilities of current audio-language models.

To construct this benchmark, the researchers developed a curated collection of 10,000 multiple-choice questions grounded in real-world audio recordings across speech, sound, and music domains. A rigorous seven-step curation and expert-annotation pipeline was established, assessing 27 distinct skills divided into information extraction and deliberate reasoning tasks. The study evaluated 18 proprietary and open-source models, as well as text-only models supplemented by automated captions, establishing a baseline comparison against human performance on a representative 1,000-question subset.

The evaluation revealed substantial performance gaps in leading artificial intelligence systems. Even the most capable multimodal models, such as Gemini Pro v1.5 and Qwen2-Audio-Instruct, achieved accuracy scores of only 52.97% and 52.50%, respectively, falling well behind the human benchmark of 82.23%. Cascaded approaches—which combine automated detailed audio captioning with high-performing text-only language models—slightly outperformed end-to-end audio models, reaching a top accuracy of 58.74%. When subjected to noise perturbation tests, several models exhibited minimal performance drops, demonstrating that they frequently rely on text-based language priors rather than genuinely attending to the audio input. Furthermore, an error analysis revealed that perceptual failures account for 55% to 64% of total mistakes, showing that basic misinterpretation of auditory signals is the primary bottleneck before complex reasoning can even occur.

These findings indicate that current audio-language models pose significant operational risks if deployed in high-stakes environments that require reliable acoustic interpretation, such as medical diagnostics, legal transcription, or autonomous monitoring. While automated captioning pipelines offer a practical interim solution, end-to-end architectures require fundamental improvements in acoustic feature extraction and cross-modal grounding to avoid reliance on superficial language patterns.

Organizations developing or deploying multimodal AI should focus future efforts on expanding training datasets with rich perceptual annotations and refining model architectures to better capture complex acoustic details. In the near term, teams should consider hybrid workflows that leverage high-quality captioning systems paired with large language models, while maintaining human oversight for critical audio analysis tasks.

The study's conclusions are constrained by its focus on English-language, multiple-choice formats and disjoint skill categorizations, which do not fully replicate open-ended conversational settings. Nevertheless, the rigorous multi-stage expert review and consistent testing methodology provide high confidence that current models face major perceptual hurdles that must be resolved before achieving expert-level audio intelligence.

  • Paper: Qwen3.5-Omni Technical Report, Qwen Team (2026). Qwen3.5-Omni directly evaluates its unified omnimodal architecture against the MMAU benchmark to demonstrate state-of-the-art multi-task audio reasoning and understanding.
  • Paper: Gemma 4 Technical Report, Gemma Team et al. (2026). Gemma 4 extends open-weight multimodal foundation models by incorporating native audio understanding and explicit thinking-mode reasoning evaluated across broad acoustic tasks.
  • Paper: Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis, Chaoyou Fu et al. (2025). Video-MME complements and extends MMAU's audio-centric evaluation by benchmarking multimodal reasoning across dynamic video contexts integrating synchronized audio tracks and subtitles.
Cover for MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark

Abstract

The ability to comprehend audio--which includes speech, non-speech sounds, and music--is crucial for AI agents to interact effectively with the world. We present MMAU, a novel benchmark designed to evaluate multimodal audio understanding models on tasks requiring expert-level knowledge and complex reasoning. MMAU comprises 10k carefully curated audio clips paired with human-annotated natural language questions and answers spanning speech, environmental sounds, and music. It includes information extraction and reasoning questions, requiring models to demonstrate 27 distinct skills across unique and challenging tasks. Unlike existing benchmarks, MMAU emphasizes advanced perception and reasoning with domain-specific knowledge, challenging models to tackle tasks akin to those faced by experts. We assess 18 open-source and proprietary (Large) Audio-Language Models, demonstrating the significant challenges posed by MMAU. Notably, even the most advanced Gemini Pro v1.5 achieves only 52.97% accuracy, and the state-of-the-art open-source Qwen2-Audio achieves only 52.50%, highlighting considerable room for improvement. We believe MMAU will drive the audio and multimodal research community to develop more advanced audio understanding models capable of solving complex audio tasks.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 The MMAU Benchmark
  • 3.1 Overview of MMAU
  • 3.2 Data Curation and Annotation
  • 3.3 Comparison with other benchmarks
  • 4 Experimental Setup
  • 5 Results and Discussion
  • 5.1 Main Results
  • 5.2 Are LALMs Really Listening?
  • 5.3 Can Captions Bridge the Gap for Text-Only Models?
  • 5.4 Deep Dive: Skill-Specific Model Performance
  • 5.5 Pinpointing LALM Weaknesses: Where Are They Falling Short?
  • 6 Conclusion, Limitations and Future Work
  • References
  • A Appendix
  • B Additional Results
  • B.1 Audio-Language Encoders (ALEs)
  • B.2 Evaluating ALEs and LALMs Across Varying Difficulty Levels
  • C Annotation Details
  • C.1 Annotation
  • C.2 Annotator Details
  • C.3 Annotation Guidelines
  • C.4 Human Evaluation
  • D Model Details
  • E Dataset Details
  • F Annotation Tool
  • G Comparison
  • H Additional Information on Skills
  • I Failure cases
  • J Benchmark Evaluation
  • K Additional Details on Error Types
  • L Prompts

Citation

MLA
Sakshi, S., et al. “MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark”. arXiv, 2024, http://arxiv.org/abs/2410.19168v1.
APA
Sakshi, S., Tyagi, U., Kumar, S., Seth, A., Selvakumar, R., Nieto, O., Duraiswami, R., Ghosh, S., & Manocha, D. (2024). MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark. arXiv. http://arxiv.org/abs/2410.19168v1
Chicago
Sakshi, S., U. Tyagi, S. Kumar, et al. 2024. “MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark”. arXiv. http://arxiv.org/abs/2410.19168v1.
Harvard
Sakshi, S. et al. (2024) “MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2410.19168v1.
Vancouver
1. Sakshi S, Tyagi U, Kumar S, Seth A, Selvakumar R, Nieto O, Duraiswami R, Ghosh S, Manocha D (2024) MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark. arXiv

BibTeX

@article{sakshi2024mmau,
  title = {MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark},
  author = {Sakshi, S and Tyagi, Utkarsh and Kumar, Sonal and Seth, Ashish and Selvakumar, Ramaneswaran and Nieto, Oriol and Duraiswami, Ramani and Ghosh, Sreyan and Manocha, Dinesh},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2410.19168v1},
  eprint = {2410.19168}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/