Audio-Reasoner: Improving Reasoning Capability in Large Audio Language Models

Zhifei XieMingbao LinZihang LiuPengcheng WuShuicheng YanChunyan Miao

article2025EMNLP135 citations

Introduces Audio-Reasoner and a 1.2-million-sample reasoning dataset to enable structured chain-of-thought processing in large audio language models, yielding substantial accuracy gains across speech, sound, and music benchmarks.

Listen

Artificial intelligence systems increasingly excel at complex reasoning, but advances have centered almost entirely on text and visual tasks while leaving audio comprehension behind. Most existing audio language models rely on simple datasets with brief labels, causing them to falter when confronted with multi-step logical questions or long-form reasoning. The article addresses this gap by introducing Audio-Reasoner, an open-source audio language model designed to perform structured, step-by-step reasoning across sound, speech, and music tasks.

To achieve this, the article outlines the creation of CoTA, a curated dataset comprising 1.2 million reasoning-rich audio samples. Using a multi-stage data synthesis and filtering pipeline powered by commercial models, the researchers transformed simple human labels from open sources into detailed descriptions, varied question-and-answer pairs, and structured chain-of-thought pathways. Audio-Reasoner, built on an 8.4-billion-parameter architecture, was trained to execute a disciplined four-step inference process: planning the approach, captioning the relevant acoustic cues, reasoning through hypotheses step by step, and summarizing the final response.

Empirical evaluations demonstrate that Audio-Reasoner establishes strong performance benchmarks across diverse audio domains. On the MMAU-mini multimodal audio reasoning benchmark, it achieved an overall accuracy of 61.71%, gaining 12.51 percentage points over its base open-source model and outperforming leading closed-source systems such as GPT-4o and Gemini-1.5-Pro. On the AIR-Bench conversational and foundational benchmarks, it posted top scores, including an average conversational evaluation of 7.94 out of 10. The system also delivered substantial performance increases in specialized domains, improving speech-to-text translation BLEU scores on CoVoST 2 by roughly 30% over its baseline and raising emotion recognition accuracy on the MELD dataset to 53.9%.

These findings indicate that incorporating structured reasoning protocols into audio models significantly improves interpretability and reduces unsupported conclusions, known as hallucinations. By breaking down analysis into explicit intermediate steps, the model becomes transparent and better suited for high-stakes operational environments such as automated customer service, medical transcription, acoustic monitoring, and multilingual translation.

Organizations evaluating audio intelligence systems should consider adopting structured chain-of-thought methodologies and investing in reasoning-dense training data rather than relying solely on raw scale or simple label matching. However, leaders should note current operational boundaries before full-scale deployment: the model is optimized for single-turn interactions and has not yet been extended to long multi-turn dialogues or multi-modal visual integrations. Furthermore, the article’s error analysis reveals that 49% of model failures stem from basic acoustic perception mistakes (such as mis-hearing audio elements under noisy conditions) and 40% from gaps in domain-specific knowledge, indicating that near-term efforts should focus on robust perceptual encoding and acoustic pre-processing.

No sufficiently relevant recommendations were found.

Cover for Audio-Reasoner: Improving Reasoning Capability in Large Audio Language Models

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Audio-Reasoner
  • 3.1 Model Training with Audio Reasoning
  • 3.2 Systematic Data Preparation
  • 3.2.1 Multistage Data Generation Pipeline
  • 3.2.2 Task Taxonomy
  • 3.3 CoTA Dataset Analysis
  • 4 Experimentation
  • 4.1 Experimental Setup
  • 4.2 Main Results
  • 4.3 Case Study
  • 4.4 Error Analysis
  • 5 Conclusion
  • 6 Limitations
  • References
  • A Prompt Details
  • A.1 Prompt of Stage 1 when Processing Data (Sample from AudioSet)
  • A.2 Prompt of Stage 2 when Processing Data (Sample from AudioSet)
  • A.3 Prompt of Stage 3 when Processing Data (Sample from AudioSet)
  • B Synthetic Data Generation Pipeline
  • B.1 Synthetic Data Introduction
  • B.2 Sample from Complex Audio Dataset
  • B.3 Sample from Multi-Speaker Dataset
  • C Further Dataset Analysis
  • C.1 More Statistical Results of CoTA
  • D Results on Latest Audio-based Reasoning Benchmark
  • E More Case Studies

Knowls

  1. Knowl 1 — Audio-Reasoner uses a four-stage audio reasoning trace

    model/method

    Audio-Reasoner takes an audio signal and a text query, then generates a structured trace followed by a user-facing response. The trace has four ordered parts: planning identifies what the query requires and sets out an approach; captioning describes audio content relevant to the query, such as speech, acoustic events, or context; reasoning combines that evidence in explicit steps; and summary consolidates the result before the final response. The model is trained to produce both reasoning and the final answer. The authors propose that this structure makes outputs easier to inspect and helps keep answers grounded; these are intended benefits of the design, not separately isolated causal findings.

  2. Knowl 2 — CoTA is generated through annotation, reasoning, and validation stages

    model/method

    CoTA is built by converting audio and simple annotations into richer question-answering examples with structured reasoning. First, a commercial model uses the audio and an existing description to produce a more detailed caption and three progressively harder questions, each with an answer; questions may be open-ended or four-option multiple choice and are meant to remain grounded in the clip. Second, a model uses the audio, caption, question, and answer to construct a chain with planning, captioning, reasoning, and summary sections, followed by a final response. Third, a reviewer model checks the audio, annotations, questions, answers, and reasoning for accuracy, coherence, suitability, and hallucinations, rejecting examples judged to contain errors. The paper reports using Google Gemini models for this data-generation process.

  3. Knowl 3 — CoTA combines broad audio coverage with purpose-built synthetic data

    data/table

    The authors describe CoTA as containing about 1.2 million reasoning-rich caption and question-answer samples across speech, music, and environmental sound. They report domain shares of 38.33% speech, 14.12% music, and 47.55% environmental sound, with 14.15% synthetic samples. The dataset composition lists the following sources and quantities: Multi-Speaker, 117.4k (12.09%, synthetic); MELD, 29.2k (3.01%); CoVoST2, 224.6k (23.13%); MusicBench, 137.1k (14.12%); AudioSet, 315.2k (32.46%); Clotho, 9.3k (0.93%); AudioCaps, 117.5k (12.10%); and Complex Audio, 20k (2.06%, synthetic). These sources support speech QA, emotion recognition, speech-to-text translation, music QA, sound QA, multi-speaker QA, and complex-audio QA. The listed synthetic sets target distinct challenges: Multi-Speaker uses generated conversations synthesized with varied speaker timbres, while Complex Audio includes temporally arranged clips for ordering or counting and blended tracks requiring identification of a target sound.

  4. Knowl 4 — Audio-Reasoner is fine-tuned to predict reasoning and answer sequences

    equation

    Audio-Reasoner is based on the 8.4-billion-parameter Qwen2-Audio-Instruct model. For each of NN training examples, let AiA_i be an audio input, QiQ_i its text query, CiC_i the structured reasoning sequence, and RiR_i the final response; let θ\theta denote the model parameters. The training objective maximizes the conditional likelihood of the reasoning and answer given the audio and query, equivalently minimizing the negative log-likelihood:

    L(θ)=−∑i=1Nlog⁡P(Ci,Ri∣Ai,Qi;θ).\mathcal{L}(\theta)=-\sum_{i=1}^{N}\log P(C_i,R_i\mid A_i,Q_i;\theta).

    The model was trained with full-parameter supervised fine-tuning using ms-swift, a peak learning rate of 10−510^{-5}, and one epoch over CoTA. The reported context window is 4K tokens, and the authors state that generated reasoning can exceed 1K tokens.

  5. Knowl 5 — Audio-Reasoner achieves the highest reported MMAU-mini average

    empirical result

    On MMAU-mini, which evaluates sound, music, and speech reasoning accuracy, Audio-Reasoner scores 60.06% on sound, 64.30% on music, and 60.70% on speech, for a 61.71% average. The Qwen2-Audio-Instruct base model scores 54.95%, 50.98%, and 42.04%, respectively, averaging 49.20%; Audio-Reasoner is therefore 12.51 percentage points higher on the average, equivalent to the paper's reported 25.42% relative gain. Its average also exceeds the reported closed-source results for GPT-4o plus captioning (57.30%) and Gemini-1.5-Pro (54.90%).

  6. Knowl 6 — Audio-Reasoner leads AIR-Bench average scores

    empirical result

    On AIR-Bench Chat, Audio-Reasoner scores 7.68 on sound, 8.05 on music, 8.19 on speech, and 6.65 on mixed audio, with an average of 7.94. The highest listed open-source baseline average is SALMONN's 6.11; Qwen-Audio-Turbo, a listed closed-source baseline, averages 6.34. On AIR-Bench Foundation, Audio-Reasoner averages 65.2, compared with 59.2 for Qwen-Audio-Turbo and 53.8 for Qwen-Audio-Chat. Its listed foundation scores are 65.7 on sound, 55.2 on music, 60.5 on Sp-SER, 88.1 on Sp-SIC, and 56.3 on Sp-SNV. The results show the highest listed average on both AIR-Bench portions, although Qwen-Audio-Turbo's music foundation score of 62.5 is higher than Audio-Reasoner's 55.2.

  7. Knowl 7 — Translation and emotion recognition improve over the base model

    empirical result

    On CoVoST 2 speech-to-text translation, Audio-Reasoner obtains an average BLEU score of 50.87 for EN-ZN and 29.13 for ZN-EN, with a reported combined average of 40.00. Qwen2-Audio-Instruct scores 37.07, 24.18, and 30.63 on the same measures; Gemini-1.5-Pro scores 46.24, 26.39, and 36.32. On the MELD speech emotion recognition test, Audio-Reasoner reaches 53.9 unweighted accuracy, compared with 49.9 for Qwen2-Audio-Instruct, 39.2 for SALMONN, and 31.5 for EMO-box.

  8. Knowl 8 — Audio-Reasoner outperforms other listed open models on MMAR

    empirical result

    On the MMAR audio-reasoning benchmark, Audio-Reasoner has a 36.80% average, the highest among the listed open-source models. The next-highest listed open-source average is 33.20% for SALAMONN 13B. Audio-Reasoner's category scores are 43.64% for sound, 33.50% for music, and 32.99% for speech; it does not lead every category, since the listed SALAMONN models score 34.69% on speech. The closed-source GPT-4o mini Audio result is 50.60% on average, above Audio-Reasoner.

  9. Knowl 9 — Most MMAU-mini errors arise from perception or knowledge failures

    data/table

    In an analysis of Audio-Reasoner's incorrect MMAU-mini answers, 49% of errors are categorized as perceptual, including mishearing and acoustic-segmentation failures, and 40% as knowledge errors, such as missing domain facts or unit conventions. The remaining errors are reasoning slips (3%), instruction misunderstanding or misreading (3%), choice-format errors (2%), and repeated thinking without an answer (3%). Thus, perception and knowledge together account for 89% of the analyzed failures; the paper identifies these as the dominant bottlenecks in the model's audio reasoning.

  10. Knowl 10 — The model's evaluation leaves important generalization limits unresolved

    limitation

    The authors state that Audio-Reasoner is primarily designed for single-turn reasoning and may have difficulty maintaining context in more complex multi-turn scenarios. Its generalization across diverse audio domains and real-world noise conditions has not been fully validated. Cross-modal reasoning, particularly combining audio with visual or textual cues, is also unexplored.

Coverage note — Qualitative case studies and detailed per-dataset response-length and audio-duration distributions are omitted because they illustrate model behavior or dataset variation but add no distinct core method or benchmark result beyond the knowls above.

References

  1. 1.Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774.
  2. 2.Cameron B Browne, Edward Powley, Daniel Whitehouse, Simon M Lucas, Peter I Cowling, Philipp Rohlfshagen, Stephen Tavener, Diego Perez, Spyridon Samothrakis, and Simon Colton. 2012. A survey of monte carlo tree search methods. IEEE Transactions on Computational Intelligence and AI in Games (T-CIAIG), (1):1–43.
  3. 3.Yunfei Chu, Jin Xu, Qian Yang, Haojie Wei, Xipin Wei, Zhifang Guo, Yichong Leng, Yuanjun Lv, Jinzheng He, Junyang Lin, et al. 2024. Qwen2-audio technical report. arXiv preprint arXiv:2407.10759.
  4. 4.Yunfei Chu, Jin Xu, Xiaohuan Zhou, Qian Yang, Shiliang Zhang, Zhijie Yan, Chang Zhou, and Jingren Zhou. 2023. Qwen-audio: Advancing universal audio understanding via unified large-scale audio-language models. arXiv preprint arXiv:2311.07919.
  5. 5.Yingqian Cui, Pengfei He, Xianfeng Tang, Qi He, Chen Luo, Jiliang Tang, and Yue Xing. 2024. A theoretical understanding of chain-of-thought: Coherent reasoning and error-aware demonstration. arXiv preprint arXiv:2410.16540.
  6. 6.Alexandre Défossez, Laurent Mazaré, Manu Orsini, Amélie Royer, Patrick Pérez, Hervé Jégou, Edouard Grave, and Neil Zeghidour. 2024. Moshi: a speech-text foundation model for real-time dialogue. arXiv preprint arXiv:2410.00037.
  7. 7.Konstantinos Drossos, Samuel Lipping, and Tuomas Virtanen. 2020. Clotho: An audio captioning dataset. In International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 736–740.
  8. 8.Zhihao Du, Yuxuan Wang, Qian Chen, Xian Shi, Xiang Lv, Tianyu Zhao, Zhifu Gao, Yexin Yang, Changfeng Gao, Hui Wang, et al. 2024. Cosyvoice 2: Scalable streaming speech synthesis with large language models. arXiv preprint arXiv:2412.10117.
  9. 9.Hao Fei, Shengqiong Wu, Wei Ji, Hanwang Zhang, Meishan Zhang, Mong Li Lee, and Wynne Hsu. 2024. Video-of-thought: step-by-step video reasoning from perception to cognition. In International Conference on Machine Learning (ICML), pages 13109–13125.
  10. 10.Chaoyou Fu, Haojia Lin, Xiong Wang, Yi-Fan Zhang, Yunhang Shen, Xiaoyu Liu, Yangze Li, Zuwei Long, Heting Gao, Ke Li, et al. 2025. Vita-1.5: Towards gpt-4o level real-time vision and speech interaction. arXiv preprint arXiv:2501.01957.
  11. 11.Jort F Gemmeke, Daniel PW Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R Channing Moore, Manoj Plakal, and Marvin Ritter. 2017. Audio set: An ontology and human-labeled dataset for audio events. In International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 776–780. IEEE.
  12. 12.Sreyan Ghosh, Sonal Kumar, Ashish Seth, Chandra Kiran Reddy Evuru, Utkarsh Tyagi, S Sakshi, Oriol Nieto, Ramani Duraiswami, and Dinesh Manocha. 2024. Gama: A large audio-language model with advanced audio understanding and complex reasoning abilities. In Empirical Methods in Natural Language Processing, pages 6288–6313.
  13. 13.Yuan Gong, Alexander H Liu, Hongyin Luo, Leonid Karlinsky, and James Glass. 2023a. Joint audio and speech understanding. In Automatic Speech Recognition and Understanding Workshop (ASRU), pages 1–8.
  14. 14.Yuan Gong, Hongyin Luo, Alexander H Liu, Leonid Karlinsky, and James Glass. 2023b. Listen, think, and understand. arXiv preprint arXiv:2305.10790.
  15. 15.Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, et al. 2025. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning. arXiv preprint arXiv:2501.12948.
  16. 16.Jarvis Guo, Tuney Zheng, Yuelin Bai, Bo Li, Yubo Wang, King Zhu, Yizhi Li, Graham Neubig, Wenhu Chen, and Xiang Yue. 2024. Mammoth-vl: Eliciting multimodal reasoning with instruction tuning at scale. arXiv preprint arXiv:2412.05237.
  17. 17.Songhao Han, Wei Huang, Hairong Shi, Le Zhuo, Xiu Su, Shifeng Zhang, Xu Zhou, Xiaojuan Qi, Yue Liao, and Si Liu. 2024. Videoespresso: A large-scale chain-of-thought dataset for fine-grained video reasoning via core frame selection. arXiv preprint arXiv:2411.14794.
  18. 18.Aaron Hurst, Adam Lerer, Adam P Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, Alan Hayes, Alec Radford, et al. 2024. Gpt-4o system card. arXiv preprint arXiv:2410.21276.
  19. 19.Aaron Jaech, Adam Kalai, Adam Lerer, Adam Richardson, Ahmed El-Kishky, Aiden Low, Alec Helyar, Aleksander Madry, Alex Beutel, Alex Carney, et al. 2024. Openai o1 system card. arXiv preprint arXiv:2412.16720.
  20. 20.Feihu Jin, Yifan Liu, and Ying Tan. 2024. Zero-shot chain-of-thought reasoning guided by evolutionary algorithms in large language models. arXiv preprint arXiv:2402.05376.
  21. 21.Chris Dongjoo Kim, Byeongchang Kim, Hyunmin Lee, and Gunhee Kim. 2019. Audiocaps: Generating captions for audios in the wild. In Nations of the Americas Chapter of the Association for Computational Linguistics (NAACL), pages 119–132.
  22. 22.Zhifeng Kong, Arushi Goel, Rohan Badlani, Wei Ping, Rafael Valle, and Bryan Catanzaro. 2024. Audio flamingo: A novel audio language model with few-shot learning and dialogue abilities. In International Conference on Machine Learning (ICML), pages 25125–25148.
  23. 23.Aixin Liu, Bei Feng, Bing Xue, Bingxuan Wang, Bochao Wu, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, et al. 2024a. Deepseek-v3 technical report. arXiv preprint arXiv:2412.19437.
  24. 24.Shansong Liu, Atin Sakkeer Hussain, Chenshuo Sun, and Ying Shan. 2024b. Music understanding llama: Advancing text-to-music generation with question answering and captioning. In International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 286–290.
  25. 25.Ziyang Ma, Mingjie Chen, Hezhao Zhang, Zhisheng Zheng, Wenxi Chen, Xiquan Li, Jiaxin Ye, Xie Chen, and Thomas Hain. 2024. Emobox: Multilingual multi-corpus speech emotion recognition toolkit and benchmark. arXiv preprint arXiv:2406.07162.
  26. 26.Ziyang Ma, Zhuo Chen, Yuping Wang, Eng Siong Chng, and Xie Chen. 2025a. Audio-cot: Exploring chain-of-thought reasoning in large audio language model. arXiv preprint arXiv:2501.07246.
  27. 27.Ziyang Ma, Yinghao Ma, Yanqiao Zhu, Chen Yang, Yi-Wen Chao, Ruiyang Xu, Wenxi Chen, Yuanzhe Chen, Zhuo Chen, Jian Cong, et al. 2025b. Mmar: A challenging benchmark for deep reasoning in speech, audio, music, and their mix. arXiv preprint arXiv:2505.13032.
  28. 28.Jan Melechovsky, Zixun Guo, Deepanway Ghosal, Navonil Majumder, Dorien Herremans, and Soujanya Poria. 2024. Mustango: Toward controllable text-to-music generation. In Nations of the Americas Chapter of the Association for Computational Linguistics (NAACL), pages 8286–8309.
  29. 29.Niklas Muennighoff, Zitong Yang, Weijia Shi, Xiang Lisa Li, Li Fei-Fei, Hannaneh Hajishirzi, Luke Zettlemoyer, Percy Liang, Emmanuel Candès, and Tatsunori Hashimoto. 2025. s1: Simple test-time scaling. arXiv preprint arXiv:2501.19393.
  30. 30.Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur. 2015. Librispeech: an asr corpus based on public domain audio books. In International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 5206–5210.
  31. 31.Soujanya Poria, Devamanyu Hazarika, Navonil Majumder, Gautam Naik, Erik Cambria, and Rada Mihalcea. 2019. Meld: A multimodal multi-party dataset for emotion recognition in conversations. In Annual Meeting of the Association for Computational Linguistics (ACL), pages 527–536.
  32. 32.Yujia Qin, Shihao Liang, Yining Ye, Kunlun Zhu, Lan Yan, Yaxi Lu, Yankai Lin, Xin Cong, Xiangru Tang, Bill Qian, et al. 2023. Toolllm: Facilitating large language models to master 16000+ real-world apis. arXiv preprint arXiv:2307.16789.
  33. 33.Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever. 2023. Robust speech recognition via large-scale weak supervision. In International Conference on Machine Learning (ICML), pages 28492–28518.
  34. 34.Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn. 2023. Direct preference optimization: Your language model is secretly a reward model. In Advances in Neural Information Processing Systems (NeurIPS), pages 53728–53741.
  35. 35.S Sakshi, Utkarsh Tyagi, Sonal Kumar, Ashish Seth, Ramaneswaran Selvakumar, Oriol Nieto, Ramani Duraiswami, Sreyan Ghosh, and Dinesh Manocha. 2024. Mmau: A massive multi-task audio understanding and reasoning benchmark. In International Conference on Learning Representations (ICLR).
  36. 36.Hao Shao, Shengju Qian, Han Xiao, Guanglu Song, Zhuofan Zong, Letian Wang, Yu Liu, and Hongsheng Li. 2024. Visual cot: Unleashing chain-of-thought reasoning in multi-modal language models. arXiv preprint arXiv:2403.16999.
  37. 37.Zayne Sprague, Fangcong Yin, Juan Diego Rodriguez, Dongwei Jiang, Manya Wadhwa, Prasann Singhal, Xinyu Zhao, Xi Ye, Kyle Mahowald, and Greg Durrett. 2024. To cot or not to cot? chain-of-thought helps mainly on math and symbolic reasoning. arXiv preprint arXiv:2409.12183.
  38. 38.Kaya Stechly, Karthik Valmeekam, and Subbarao Kambhampati. 2024. Chain of thoughtlessness? an analysis of cot in planning. In Advances in Neural Information Processing Systems (NeurIPS), pages 29106–29141.
  39. 39.Yixuan Su, Tian Lan, Huayang Li, Jialu Xu, Yan Wang, and Deng Cai. 2023. Pandagpt: One model to instruction-follow them all. In Workshop on Taming Large Language Models: Controllability in the era of Interactive Assistants (TLLM), pages 11–23.
  40. 40.Changli Tang, Wenyi Yu, Guangzhi Sun, Xianzhao Chen, Tian Tan, Wei Li, Lu Lu, Zejun Ma, and Chao Zhang. 2023. Salmonn: Towards generic hearing abilities for large language models. arXiv preprint arXiv:2310.13289.
  41. 41.Yunlong Tang, Gen Zhan, Li Yang, Yiting Liao, and Chenliang Xu. 2024. Cardiff: Video salient object ranking chain of thought reasoning for saliency prediction with diffusion. arXiv preprint arXiv:2408.12009.
  42. 42.Gemini Team, Petko Georgiev, Ving Ian Lei, Ryan Burnell, Libin Bai, Anmol Gulati, Garrett Tanzer, Damien Vincent, Zhufeng Pan, Shibo Wang, et al. 2024. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context. arXiv preprint arXiv:2403.05530.
  43. 43.Kimi Team, Angang Du, Bofei Gao, Bowei Xing, Changjiu Jiang, Cheng Chen, Cheng Li, Chenjun Xiao, Chenzhuang Du, Chonghua Liao, et al. 2025. Kimi k1. 5: Scaling reinforcement learning with llms. arXiv preprint arXiv:2501.12599.
  44. 44.Miles Turpin, Julian Michael, Ethan Perez, and Samuel Bowman. 2023. Language models don’t always say what they think: Unfaithful explanations in chain-of-thought prompting. In Advances in Neural Information Processing Systems (NeurIPS), pages 74952–74965.
  45. 45.Changhan Wang, Anne Wu, Jiatao Gu, and Juan Pino. 2021. Covost 2 and massively multilingual speech translation. In Conference of the International Speech Communication Association (Interspeech), pages 2247–2251.
  46. 46.Chen Wang, Minpeng Liao, Zhongqiang Huang, Jinliang Lu, Junhong Wu, Yuchen Liu, Chengqing Zong, and Jiajun Zhang. 2023. Blsp: Bootstrapping language-speech pre-training via behavior alignment of continuation writing. arXiv preprint arXiv:2309.00916.
  47. 47.Yan Wang, Yawen Zeng, Jingsheng Zheng, Xiaofen Xing, Jin Xu, and Xiangmin Xu. 2024. Videocot: A video chain-of-thought dataset with active annotation tool. In Workshop on Advances in Language and Vision Research (ALVR), pages 92–101.
  48. 48.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022. Chain-of-thought prompting elicits reasoning in large language models. In Advances in Neural Information Processing Systems (NeurIPS), pages 24824–24837.
  49. 49.Peng Wen, Teng-Gen Hu, Robert J Linhardt, Sen-Tai Liao, Hong Wu, and Yu-Xiao Zou. 2019. Mulberry: A review of bioactive compounds and advanced processing technology. Trends in food science & technology, 83:138–158.
  50. 50.Shengqiong Wu, Hao Fei, Leigang Qu, Wei Ji, and Tat-Seng Chua. 2024. Next-gpt: Any-to-any multimodal llm. In International Conference on Machine Learning (ICML), pages 53366–53397.
  51. 51.Zhifei Xie and Changqiao Wu. 2024a. Mini-omni: Language models can hear, talk while thinking in streaming. arXiv preprint arXiv:2408.16725.
  52. 52.Zhifei Xie and Changqiao Wu. 2024b. Mini-omni2: Towards open-source gpt-4o with vision, speech and duplex capabilities. arXiv preprint arXiv:2410.11190.
  53. 53.Guowei Xu, Peng Jin, Li Hao, Yibing Song, Lichao Sun, and Li Yuan. 2024. Llava-o1: Let vision language models reason step-by-step. arXiv preprint arXiv:2411.10440.
  54. 54.An Yang, Beichen Zhang, Binyuan Hui, Bofei Gao, Bowen Yu, Chengpeng Li, Dayiheng Liu, Jianhong Tu, Jingren Zhou, Junyang Lin, et al. 2024a. Qwen2.5-math technical report: Toward mathematical expert model via self-improvement. arXiv preprint arXiv:2409.12122.
  55. 55.Qian Yang, Jin Xu, Wenrui Liu, Yunfei Chu, Ziyue Jiang, Xiaohuan Zhou, Yichong Leng, Yuanjun Lv, Zhou Zhao, Chang Zhou, et al. 2024b. Air-bench: Benchmarking large audio-language models via generative comprehension. In Annual Meeting of the Association for Computational Linguistics (ACL), pages 1979–1998.
  56. 56.Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Tom Griffiths, Yuan Cao, and Karthik Narasimhan. 2023. Tree of thoughts: Deliberate problem solving with large language models. In Advances in Neural Information Processing Systems (NeurIPS), pages 11809–11822.
  57. 57.Dong Zhang, Shimin Li, Xin Zhang, Jun Zhan, Pengyu Wang, Yaqian Zhou, and Xipeng Qiu. 2023a. Speechgpt: Empowering large language models with intrinsic cross-modal conversational abilities. In Empirical Methods in Natural Language Processing (EMNLP), pages 15757–15773.
  58. 58.Dong Zhang, Shimin Li, Xin Zhang, Jun Zhan, Pengyu Wang, Yaqian Zhou, and Xipeng Qiu. 2023b. Speechgpt: Empowering large language models with intrinsic cross-modal conversational abilities. arXiv preprint arXiv:2305.11000.
  59. 59.Ruohong Zhang, Bowen Zhang, Yanghao Li, Haotian Zhang, Zhiqing Sun, Zhe Gan, Yinfei Yang, Ruoming Pang, and Yiming Yang. 2024a. Improve vision language model chain-of-thought reasoning. arXiv preprint arXiv:2410.16198.
  60. 60.Yuxiang Zhang, Shangxi Wu, Yuqi Yang, Jiangming Shu, Jinlin Xiao, Chao Kong, and Jitao Sang. 2024b. o1-coder: an o1 replication for coding. arXiv preprint arXiv:2412.00154.
  61. 61.Zhuosheng Zhang, Aston Zhang, Mu Li, and Alex Smola. 2022. Automatic chain of thought prompting in large language models. arXiv preprint arXiv:2210.03493.
  62. 62.Yu Zhao, Huifeng Yin, Bo Zeng, Hao Wang, Tianqi Shi, Chenyang Lyu, Longyue Wang, Weihua Luo, and Kaifu Zhang. 2024a. Marco-o1: Towards open reasoning models for open-ended solutions. arXiv preprint arXiv:2411.14405.
  63. 63.Yuze Zhao, Jintao Huang, Jinghan Hu, Xingjun Wang, Yunlin Mao, Daoze Zhang, Yezinzi Jiang, Zhikai Wu, Baole Ai, Ang Wang, et al. 2024b. Swift: a scalable lightweight infrastructure for fine-tuning. arXiv preprint arXiv:2408.05517.
  64. 64.Qiji Zhou, Ruochen Zhou, Zike Hu, Panzhong Lu, Siyang Gao, and Yue Zhang. 2024. Image-of-thought prompting for visual reasoning refinement in multimodal large language models. arXiv preprint arXiv:2405.13872.
  65. 65.Anni Zou, Zhuosheng Zhang, Hai Zhao, and Xiangru Tang. 2023. Generalizable chain-of-thought prompting in mixed-task scenarios with large language models. arXiv preprint arXiv:2310.06692.

Citation

MLA
Zhifei, X., et al. “Audio-Reasoner: Improving Reasoning Capability in Large Audio Language Models”. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, 2025, pp. 23829–51, https://doi.org/10.18653/v1/2025.emnlp-main.1216.
APA
Zhifei, X., Lin, M., Liu, Z., Wu, P., Yan, S., & Miao, C. (2025). Audio-Reasoner: Improving Reasoning Capability in Large Audio Language Models. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, 23829–23851. https://doi.org/10.18653/v1/2025.emnlp-main.1216
Chicago
Zhifei, X., M. Lin, Z. Liu, P. Wu, S. Yan, and C. Miao. 2025. “Audio-Reasoner: Improving Reasoning Capability in Large Audio Language Models”. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, 23829–51. https://doi.org/10.18653/v1/2025.emnlp-main.1216.
Harvard
Zhifei, X. et al. (2025) “Audio-Reasoner: Improving Reasoning Capability in Large Audio Language Models”, Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 23829–23851. Available at: https://doi.org/10.18653/v1/2025.emnlp-main.1216.
Vancouver
1. Zhifei X, Lin M, Liu Z, Wu P, Yan S, Miao C (2025) Audio-Reasoner: Improving Reasoning Capability in Large Audio Language Models. In: Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 23829–23851

BibTeX

@inproceedings{zhifei-etal-2025-audio,
    title = "Audio-Reasoner: Improving Reasoning Capability in Large Audio Language Models",
    author = "Zhifei, Xie  and
      Lin, Mingbao  and
      Liu, Zihang  and
      Wu, Pengcheng  and
      Yan, Shuicheng  and
      Miao, Chunyan",
    editor = "Christodoulopoulos, Christos  and
      Chakraborty, Tanmoy  and
      Rose, Carolyn  and
      Peng, Violet",
    booktitle = "Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing",
    month = nov,
    year = "2025",
    address = "Suzhou, China",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2025.emnlp-main.1216/",
    doi = "10.18653/v1/2025.emnlp-main.1216",
    pages = "23829--23851",
    ISBN = "979-8-89176-332-6"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/