Audio-Reasoner: Improving Reasoning Capability in Large Audio Language Models
Zhifei XieMingbao LinZihang LiuPengcheng WuShuicheng YanChunyan Miao
Introduces Audio-Reasoner and a 1.2-million-sample reasoning dataset to enable structured chain-of-thought processing in large audio language models, yielding substantial accuracy gains across speech, sound, and music benchmarks.
Artificial intelligence systems increasingly excel at complex reasoning, but advances have centered almost entirely on text and visual tasks while leaving audio comprehension behind. Most existing audio language models rely on simple datasets with brief labels, causing them to falter when confronted with multi-step logical questions or long-form reasoning. The article addresses this gap by introducing Audio-Reasoner, an open-source audio language model designed to perform structured, step-by-step reasoning across sound, speech, and music tasks.
To achieve this, the article outlines the creation of CoTA, a curated dataset comprising 1.2 million reasoning-rich audio samples. Using a multi-stage data synthesis and filtering pipeline powered by commercial models, the researchers transformed simple human labels from open sources into detailed descriptions, varied question-and-answer pairs, and structured chain-of-thought pathways. Audio-Reasoner, built on an 8.4-billion-parameter architecture, was trained to execute a disciplined four-step inference process: planning the approach, captioning the relevant acoustic cues, reasoning through hypotheses step by step, and summarizing the final response.
Empirical evaluations demonstrate that Audio-Reasoner establishes strong performance benchmarks across diverse audio domains. On the MMAU-mini multimodal audio reasoning benchmark, it achieved an overall accuracy of 61.71%, gaining 12.51 percentage points over its base open-source model and outperforming leading closed-source systems such as GPT-4o and Gemini-1.5-Pro. On the AIR-Bench conversational and foundational benchmarks, it posted top scores, including an average conversational evaluation of 7.94 out of 10. The system also delivered substantial performance increases in specialized domains, improving speech-to-text translation BLEU scores on CoVoST 2 by roughly 30% over its baseline and raising emotion recognition accuracy on the MELD dataset to 53.9%.
These findings indicate that incorporating structured reasoning protocols into audio models significantly improves interpretability and reduces unsupported conclusions, known as hallucinations. By breaking down analysis into explicit intermediate steps, the model becomes transparent and better suited for high-stakes operational environments such as automated customer service, medical transcription, acoustic monitoring, and multilingual translation.
Organizations evaluating audio intelligence systems should consider adopting structured chain-of-thought methodologies and investing in reasoning-dense training data rather than relying solely on raw scale or simple label matching. However, leaders should note current operational boundaries before full-scale deployment: the model is optimized for single-turn interactions and has not yet been extended to long multi-turn dialogues or multi-modal visual integrations. Furthermore, the article’s error analysis reveals that 49% of model failures stem from basic acoustic perception mistakes (such as mis-hearing audio elements under noisy conditions) and 40% from gaps in domain-specific knowledge, indicating that near-term efforts should focus on robust perceptual encoding and acoustic pre-processing.
- Paper: MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark, S. Sakshi et al. (2025). Read this benchmark first to understand the MMAU-mini evaluation that Audio-Reasoner uses to measure audio reasoning performance.
- Paper: AIR-Bench: Benchmarking Large Audio-Language Models via Generative Comprehension, Qian Yang et al. (2024). Its AIR-Bench evaluation framework provides the context needed to interpret Audio-Reasoner’s reported conversational and foundational benchmark results.
No sufficiently relevant recommendations were found.
