AudioChatLlama: Towards General-Purpose Speech Abilities for LLMs

Yassir FathullahChunyang WuEgor LakomkinKe LiJunteng JiaYuan ShangguanJay MahadeokarOzlem KalinliChristian FuegenMike Seltzer

article2024NAACL84 citations

Presents AudioChatLlama, an end-to-end framework that extends instruction-tuned language models to general-purpose speech reasoning and conversational tasks without requiring paired multimodal datasets by training an audio encoder to align speech prompts with text-based responses.

Listen

Conversational artificial intelligence currently faces significant friction when interacting through voice. Most voice-enabled language systems rely on a cascaded structure that first transcribes spoken audio into text and then feeds that text into a language model. This sequential pipeline introduces compounding transcription errors, increases processing latency, and struggles to retain conversational context when encountering uncommon words or ambiguous audio.

The article demonstrates an end-to-end multi-modal framework called AudioChatLlama, which equips a conversational large language model with direct speech understanding and reasoning capabilities without requiring curated paired speech-instruction datasets.

To achieve this, the authors aligned a lightweight audio encoder to a frozen conversational text model (Llama-2-chat 7B). Rather than collecting expensive human-annotated speech-and-response datasets, the training utilized 45,000 hours of public English audio from the Multilingual LibriSpeech corpus. The authors established a principle of modal invariance: spoken inputs and their exact text transcriptions should produce identical outputs from the language model. By feeding transcripts into the language model to automatically generate responses and then training only the audio encoder to map spoken inputs to those same outputs, the system learned to process audio prompts as direct substitutes for text while leaving the base language model completely unaltered.

Evaluations across synthetic and real-world spoken question-answering benchmarks revealed three primary findings. First, AudioChatLlama achieved lower perplexity scores than cascaded systems, indicating a more accurate capture of intended responses. Second, in human evaluations of spoken questions with high speech-recognition error rates (exceeding 30%), the end-to-end model outperformed the traditional cascade by 6 to 12 percentage points in successful response rates (achieving 48–52% success compared to 40–42% for the baseline). In low-error settings, it performed on par with cascaded setups at approximately 78–80% success. Third, the system demonstrated cross-modal flexibility, successfully handling tasks such as speech translation, audio summarization, contextual multi-turn dialogue, and seamless switching between voice and text inputs within the same conversation.

These findings suggest that organizations can expand language models into voice interfaces without the high cost of training full foundation models from scratch or curating niche paired datasets. Operating end-to-end reduces failure points caused by transcription mistakes and allows conversational context to resolve spoken ambiguities naturally. However, the system's overall intelligence remains bounded by the underlying language model's capabilities, meaning weaknesses in the base text model directly persist in the speech domain.

Decision-makers exploring voice-driven artificial intelligence should consider end-to-end alignment strategies over traditional transcription cascades, particularly for workflows involving complex terminology or noisy acoustic environments. Before production deployment, development teams should pilot larger self-supervised audio encoders and refine speech segmentation to ensure higher training data quality. While confidence is high regarding the model's robustness over standard cascaded pipelines, caution is warranted around phonetically similar words and proper names, where the system can still occasionally confuse distinct acoustic entities.

Fathullah et al (2024).pdf
Cover for AudioChatLlama: Towards General-Purpose Speech Abilities for LLMs

Abstract

In this work, we extend the instruction-tuned Llama-2 model with end-to-end general-purpose speech processing and reasoning abilities while maintaining the wide range of original LLM capabilities, without using any carefully curated paired data. The resulting end-to-end model, named AudioChatLlama, can utilize audio prompts as a replacement for text and sustain a conversation. Such a model also has extended cross-modal capabilities such as being able to perform spoken question answering (QA), speech translation, and audio summarization amongst many other closed and open-domain tasks. This is unlike prior approaches in speech, in which LLMs are extended to handle audio for a limited number of pre-designated tasks. On both synthesized and recorded speech QA test sets, evaluations show that our end-to-end approach is on par with or outperforms cascaded systems (speech recognizer + LLM) in terms of modelling the response to a prompt. Furthermore, unlike cascades, our approach can interchange text and audio modalities and intrinsically utilize prior context in a conversation to provide better results.

Table of Contents

  • Abstract
  • 1 Introduction
  • 2 AudioChatLlama: LLM with General-Purpose Speech Abilities
  • 2.1 Architecture
  • 2.2 General-Purpose Alignment Using Unpaired Data
  • 3 Experimental Setup
  • 3.1 Dataset & Generation
  • 3.2 Model Architecture & Training
  • 4 Results
  • 4.1 Cascade Baseline
  • 4.2 Perplexity
  • 4.3 Human Evaluation
  • 4.4 Embedding Space Alignment
  • 4.5 Extended Capabilities
  • 5 Limitation
  • 6 Conclusion
  • References
  • A Examples of Extended Capabilities

Knowls

  1. Knowl 1 — Modal-invariant speech prompting for a conversational LLM

    model/method

    AudioChatLlama extends the instruction-tuned, conversational Llama-2-chat 7B model so that speech can serve as a direct alternative to text input. Its alignment principle is modal invariance: when speech and text convey the same semantic content, the model should produce the same response. This lets the system apply the language model’s general response-generation behavior to speech rather than restricting speech input to a fixed set of predefined tasks.

  2. Knowl 2 — Unpaired ASR data turned into speech-response training examples

    model/method

    The training method uses existing automatic speech recognition (ASR) pairs of audio and transcript without requiring curated speech-question-and-answer data. Each transcript is submitted as the user prompt to Llama-2-chat 7B, with an empty system prompt, and the generated reply becomes the target for the corresponding audio. The audio encoder is trained on these audio–reply pairs to elicit the response that the transcript elicited from the text model. The workflow diagram on page 3 depicts this transcript-to-response and audio-to-response alignment.

  3. Knowl 3 — A shared embedding-sequence interface for audio and text

    model/method

    AudioChatLlama feeds a sequence of embeddings to a decoder-only language model, which generates a text response autoregressively. For speech, variable-length audio embeddings are placed between text-embedded prompt tokens, including the instruction and any conversation history; for text input, text embeddings replace the audio encoder output. The architecture schematic on page 2 shows this common interface. Llama-2-chat 7B is kept frozen, and only the audio encoder is trained, so the extension does not update the language-model parameters.

  4. Knowl 4 — Conformer audio encoder and its alignment training

    experimental setup

    The audio encoder consumes 80-dimensional filterbank features at a 10 ms frame rate. A convolutional feature extractor reduces the rate to 80 ms, then a linear layer maps features to 512 dimensions before a stack of either 18 or 36 Conformer layers. Each layer has a 512-dimensional hidden representation, a 2048-dimensional feed-forward network, kernel size 11, and 8 attention heads. For encoder pretraining, a CTC output layer uses a 1.5k-token SentencePiece vocabulary and is discarded afterward. The 512-dimensional encoder outputs are grouped in sets of nn consecutive frames and projected to 4096 dimensions to match Llama-2-chat 7B; the resulting audio-embedding rate is 80n80n ms.

    The encoder is pretrained with Adam (β1=0.9\beta_1=0.9, β2=0.98\beta_2=0.98), using 20k warmup steps to a peak learning rate of 10−310^{-3} followed by exponential decay. Pretraining uses 16 NVIDIA A100 40 GB GPUs, 4 gradient-accumulation steps, and up to 500 seconds of audio per GPU batch. Joint audio-encoder/LLM training uses 5k warmup steps to a peak learning rate of 5×10−45\times10^{-4}, decaying to 5×10−65\times10^{-6} over 250k steps; runs were often stopped within 100k steps. It uses 64 A100 40 GB GPUs, 8 gradient-accumulation steps, and batch size 2. Response decoding uses beam search with beam size 10.

  5. Knowl 5 — Response perplexity is lower than the tested ASR cascades

    empirical result

    The evaluation measures perplexity (PPL) of each system on the response generated by Llama-2-chat from the reference transcript; lower PPL indicates closer modeling of that target response. On the MLS test set, the reference-text prompt has PPL 1.383. Three cascade ASR+LLM settings have prompt word error rates (WERs) and response PPLs of 16.8% and 1.831, 10.1% and 1.641, and 7.5% and 1.575. AudioChatLlama obtains PPL 1.559 with the 18-layer Conformer and 1.544 with the 36-layer Conformer. On TriviaQA questions rendered into speech with text-to-speech (TriviaQA-TTS), the reference-text prompt has PPL 1.273; cascade settings with WERs of 15.2%, 11.5%, and 10.3% obtain PPLs of 1.775, 1.720, and 1.709, while the 18- and 36-layer AudioChatLlama models obtain 1.467 and 1.422. Thus both end-to-end models have lower response PPL than every reported cascade setting on both test sets, though neither matches the text-prompt reference. These comparisons are reported on page 5.

  6. Knowl 6 — Human judgments favor end-to-end responses at higher ASR error rates

    empirical result

    Human evaluation compares response success rate (SR)—the fraction of responses judged to agree with the reference answer—for 50 questions at each ASR prompt-WER level. On TriviaQA-TTS, the cascade versus AudioChatLlama SRs are 40% versus 52% at 37.5% WER, 60% versus 70% at 14.3% WER, and 80% versus 80% at 4.3% WER. On the recorded RealSQA evaluation set, the corresponding results are 42% versus 48% at 31.5% WER, 50% versus 58% at 13.3% WER, and 80% versus 78% at 3.5% WER. Across these two evaluations, AudioChatLlama scores higher in the two higher-error conditions and is approximately on par with the cascade in the lowest-error condition.

  7. Knowl 7 — Evaluation data and spoken-question test sets

    experimental setup

    The experiments use the English portion of Multilingual LibriSpeech (MLS), an audiobook ASR corpus described as 50k hours overall, with 45k hours in English. Audio is segmented into utterances of up to 20 seconds. TriviaQA-TTS is made by rendering text-based TriviaQA questions into speech. RealSQA is a small recorded spoken-question evaluation set: scripts sourced from OpenAssistant are filtered with a BERT-based named-entity tagger to select 8k utterances containing at least one entity and shorter than 100 characters; 200 participants receive random script sets, with at most 200 utterances per participant. Its reported evaluation uses three ASR-error strata with 50 examples in each.

  8. Knowl 8 — Speech input supports inherited language tasks and mixed-modality dialogue

    empirical result

    Qualitative examples show AudioChatLlama answering spoken questions, summarizing spoken content, translating speech into text, and accepting text and audio on alternating turns in one conversation. The examples also illustrate contextualization: conversation history can help interpret acoustically difficult or rare names, such as Jökulsárlón in a discussion of Iceland. These capabilities are not separately trained task heads; they arise from extending the instruction-tuned LLM’s behavior through audio input, so their quality depends on what the underlying LLM can do. The paper presents demonstrations rather than a systematic quantitative evaluation of these tasks.

  9. Knowl 9 — Audio and transcript embeddings show an approximately monotonic alignment

    empirical result

    For one human-recorded example, the audio encoder produces 40 embeddings from approximately 3.2 seconds of audio at an 80 ms frame rate, while the transcript produces 10 text embeddings. The pairwise cosine-similarity heatmap on page 7 shows an approximately monotonic correspondence between audio and text embeddings despite the different sequence lengths. The alignment is weak for Obama, and the uninformative beginning and ending regions correspond to deliberate silence in the recording. This is a single-example observation supporting the authors’ expectation that the audio encoder learns representations compatible with the frozen LLM’s text-embedding space.

  10. Knowl 10 — Limited generic audio understanding and sensitivity to similar-sounding names

    limitation

    AudioChatLlama is shown to extend speech interaction and speech-related tasks, but the paper identifies its general audio understanding and reasoning as limited. Its small Conformer encoder is proposed as a starting point rather than a comprehensive audio-understanding solution; stronger self-supervised encoders are suggested as a possible next step. The model can also misidentify acoustically similar names: in one example with a cascade prompt WER of 0%, AudioChatLlama answered about Joseph Schumpeter when the question concerned Joseph Scaliger, while the cascade answered correctly. The authors also note that short, sentence-fragmented audiobook utterances can produce nonsensical prompts and that greedy generation of training replies, chosen for computational reasons, may limit target quality.

Coverage note — The appendix’s individual dialogue transcripts are not reproduced; their distinct contributions are summarized as qualitative demonstrations of translation, summarization, modality switching, and contextualization.

References

  1. 1.Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli. 2020. wav2vec 2.0: A framework for self-supervised learning of speech representations.
  2. 2.Emily M. Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021. On the dangers of stochastic parrots: Can language models be too big? In ACM Conference on Fairness, Accountability, and Transparency.
  3. 3.Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, et al. 2021. On the opportunities and risks of foundation models. arXiv preprint arXiv:2108.07258.
  4. 4.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems.
  5. 5.Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. 2023. Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.
  6. 6.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. 2022. Palm: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311.
  7. 7.Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei. 2017. Deep reinforcement learning from human preferences. In Advances in Neural Information Processing Systems.
  8. 8.Mostafa Dehghani, Josip Djolonga, Basil Mustafa, Piotr Padlewski, Jonathan Heek, Justin Gilmer, Andreas Steiner, Mathilde Caron, Robert Geirhos, Ibrahim Alabdulmohsin, et al. 2023. Scaling vision transformers to 22 billion parameters. arXiv preprint arXiv:2302.05442.
  9. 9.Danny Driess, Fei Xia, Mehdi SM Sajjadi, Corey Lynch, Aakanksha Chowdhery, Brian Ichter, Ayzaan Wahid, Jonathan Tompson, Quan Vuong, Tianhe Yu, et al. 2023. Palm-e: An embodied multimodal language model. arXiv preprint arXiv:2303.03378.
  10. 10.Yassir Fathullah, Chunyang Wu, Egor Lakomkin, Junteng Jia, Yuan Shangguan, Ke Li, Jinxi Guo, Wenhan Xiong, Jay Mahadeokar, Ozlem Kalinli, Christian Fuegen, and Mike Seltzer. 2023. Prompting large language models with speech recognition abilities. arXiv preprint arXiv:2307.11795.
  11. 11.Yuan Gong, Hongyin Luo, Alexander H Liu, Leonid Karlinsky, and James Glass. 2023. Listen, think, and understand. arXiv preprint arXiv:2305.10790.
  12. 12.Alex Graves, Santiago Fernández, Faustino Gomez, and Jürgen Schmidhuber. 2006. Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks. In Proceedings of the 23rd international conference on Machine learning, pages 369–376.
  13. 13.Anmol Gulati, James Qin, Chung-Cheng Chiu, Niki Parmar, Yu Zhang, Jiahui Yu, Wei Han, Shibo Wang, Zhengdong Zhang, Yonghui Wu, et al. 2020. Conformer: Convolution-augmented transformer for speech recognition. In Interspeech.
  14. 14.Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed. 2021. Hubert: Self-supervised speech representation learning by masked prediction of hidden units. In IEEE/ACM Transactions on Audio, Speech, and Language Processing.
  15. 15.Mandar Joshi, Eunsol Choi, Daniel S Weld, and Luke Zettlemoyer. 2017. Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension. arXiv preprint arXiv:1705.03551.
  16. 16.Zachary Kenton, Tom Everitt, Laura Weidinger, Iason Gabriel, Vladimir Mikulik, and Geoffrey Irving. 2021. Alignment of language agents. arXiv preprint arXiv:2103.14659.
  17. 17.Diederik P. Kingma and Jimmy Ba. 2015. Adam: A method for stochastic optimization. In International Conference on Learning Representations (ICLR).
  18. 18.Taku Kudo and John Richardson. 2018. SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing: System Demonstrations.
  19. 19.Egor Lakomkin, Chunyang Wu, Yassir Fathullah, Ozlem Kalinli, Michael L. Seltzer, and Christian Fuegen. 2023. End-to-end speech recognition contextualization with large language models. arXiv preprint arXiv:2309.10917.
  20. 20.Jan Leike, David Krueger, Tom Everitt, Miljan Martic, Vishal Maini, and Shane Legg. 2018. Scalable agent alignment via reward modeling: a research direction. arXiv preprint arXiv:1811.07871.
  21. 21.Shuo Liu, Leda Sarı, Chunyang Wu, Gil Keren, Yuan Shangguan, Jay Mahadeokar, and Ozlem Kalinli. 2023. Towards selection of text-to-speech data to augment asr training. arXiv preprint arXiv:2306.00998.
  22. 22.Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli. 2019. fairseq: A fast, extensible toolkit for sequence modeling. arXiv preprint arXiv:1904.01038.
  23. 23.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Gray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and Ryan Lowe. 2022. Training language models to follow instructions with human feedback. In Advances in Neural Information Processing Systems.
  24. 24.Vineel Pratap, Qiantong Xu, Anuroop Sriram, Gabriel Synnaeve, and Ronan Collobert. 2020. MLS: A large-scale multilingual dataset for speech research. In Interspeech.
  25. 25.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019. Language models are unsupervised multitask learners. OpenAI blog.
  26. 26.Andrew Rosenberg, Yu Zhang, Bhuvana Ramabhadran, Ye Jia, Pedro Moreno, Yonghui Wu, and Zelin Wu. 2019. Speech recognition with augmented synthesized speech. In IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), pages 996–1002.
  27. 27.Nick Rossenbach, Albert Zeyer, Ralf Schlüter, and Hermann Ney. 2020. Generating synthetic audio data for attention-based speech recognition systems. In IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 7069–7073.
  28. 28.Paul K Rubenstein, Chulayuth Asawaroengchai, Duc Dung Nguyen, Ankur Bapna, Zalán Borsos, Félix de Chaumont Quitry, Peter Chen, Dalia El Badawy, Wei Han, Eugene Kharitonov, et al. 2023. Audiopalm: A large language model that can speak and listen. arXiv preprint arXiv:2306.12925.
  29. 29.Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilic, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, Matthias Gallé, et al. 2022. Bloom: A 176b-parameter open-access multilingual language model. arXiv preprint arXiv:2211.05100.
  30. 30.Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F Christiano. 2020. Learning to summarize with human feedback.
  31. 31.Changli Tang, Wenyi Yu, Guangzhi Sun, Xianzhao Chen, Tian Tan, Wei Li, Lu Lu, Zejun Ma, and Chao Zhang. 2023. Salmonn: Towards generic hearing abilities for large language models. arXiv preprint arXiv:2310.13289.
  32. 32.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023a. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971.
  33. 33.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, et al. 2023b. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288.
  34. 34.Laura Weidinger, John Mellor, Maribeth Rauh, Conor Griffin, Jonathan Uesato, Po-Sen Huang, Myra Cheng, Mia Glaese, Borja Balle, Atoosa Kasirzadeh, et al. 2021. Ethical and social risks of harm from language models. arXiv preprint arXiv:2112.04359.
  35. 35.Jian Wu, Yashesh Gaur, Zhuo Chen, Long Zhou, Yimeng Zhu, Tianrui Wang, Jinyu Li, Shujie Liu, Bo Ren, Linquan Liu, and Yu Wu. 2023. On decoder-only architecture for speech-to-text and large language model integration. arXiv preprint arXiv:2307.03917.
  36. 36.Wenyi Yu, Changli Tang, Guangzhi Sun, Xianzhao Chen, Tian Tan, Wei Li, Lu Lu, Zejun Ma, and Chao Zhang. 2023. Connecting speech encoder and large language model for asr. arXiv preprint arXiv:2309.13963.
  37. 37.Xianrui Zheng, Yulan Liu, Deniz Gunceler, and Daniel Willett. 2021. Using synthetic audio to improve the recognition of out-of-vocabulary words in end-to-end asr systems. In ICASSP 2021 - 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 5674–5678.
  38. 38.Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny. 2023. Minigpt-4: Enhancing vision-language understanding with advanced large language models. arXiv preprint arXiv:2304.10592.

Citation

MLA
Fathullah, Y., et al. “AudioChatLlama: Towards General-Purpose Speech Abilities for LLMs”. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 2024, pp. 5522–32, https://doi.org/10.18653/v1/2024.naacl-long.309.
APA
Fathullah, Y., Wu, C., Lakomkin, E., Li, K., Jia, J., Shangguan, Y., Mahadeokar, J., Kalinli, O., Fuegen, C., & Seltzer, M. (2024). AudioChatLlama: Towards General-Purpose Speech Abilities for LLMs. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 5522–5532. https://doi.org/10.18653/v1/2024.naacl-long.309
Chicago
Fathullah, Y., C. Wu, E. Lakomkin, et al. 2024. “AudioChatLlama: Towards General-Purpose Speech Abilities for LLMs”. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 5522–32. https://doi.org/10.18653/v1/2024.naacl-long.309.
Harvard
Fathullah, Y. et al. (2024) “AudioChatLlama: Towards General-Purpose Speech Abilities for LLMs”, Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, pp. 5522–5532. Available at: https://doi.org/10.18653/v1/2024.naacl-long.309.
Vancouver
1. Fathullah Y, Wu C, Lakomkin E, Li K, Jia J, Shangguan Y, Mahadeokar J, Kalinli O, Fuegen C, Seltzer M (2024) AudioChatLlama: Towards General-Purpose Speech Abilities for LLMs. In: Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, pp 5522–5532

BibTeX

@inproceedings{fathullah-etal-2024-audiochatllama,
    title = "{A}udio{C}hat{L}lama: Towards General-Purpose Speech Abilities for {LLM}s",
    author = "Fathullah, Yassir  and
      Wu, Chunyang  and
      Lakomkin, Egor  and
      Li, Ke  and
      Jia, Junteng  and
      Shangguan, Yuan  and
      Mahadeokar, Jay  and
      Kalinli, Ozlem  and
      Fuegen, Christian  and
      Seltzer, Mike",
    editor = "Duh, Kevin  and
      Gomez, Helena  and
      Bethard, Steven",
    booktitle = "Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)",
    month = jun,
    year = "2024",
    address = "Mexico City, Mexico",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.naacl-long.309/",
    doi = "10.18653/v1/2024.naacl-long.309",
    pages = "5522--5532"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/