Advancing Large Language Models to Capture Varied Speaking Styles and Respond Properly in Spoken Conversations

Guan-Ting LinCheng-Han ChiangHung-yi Lee

article2024ACL75 citations

Introduces the StyleTalk benchmark and Spoken-LLM framework to enable language models to interpret prosodic variations in spoken inputs and generate style-appropriate spoken dialogue responses.

Listen

In spoken human interaction, identical words convey drastically different meanings depending on how they are delivered. Factors such as tone, emotion, speed, and volume fundamentally alter conversational intent, yet text-only large language models fail to perceive these vocal nuances. Current conversational systems typically convert speech into flat text transcripts, discarding essential paralinguistic cues and causing automated agents to deliver misaligned or unnatural spoken responses.

The article evaluates whether multimodal language models can be trained to recognize varied speaking styles from audio input and generate contextually appropriate response texts alongside matching expressive vocal styles. To accomplish this, the authors investigate model architectures and data pipelines capable of handling conversational scenarios where identical spoken phrases require divergent responses based entirely on vocal delivery.

To address the absence of relevant benchmarks, the researchers created StyleTalk, a speech-to-speech dataset spanning 17 everyday topics generated via automated language models, synthesized using expressive text-to-speech tools, and filtered by human evaluators. They then introduced Spoken-LLM, a multimodal framework integrating a frozen speech emotion encoder, a parameter-efficient language model adapter, and an expressive voice synthesizer. The system uses a two-stage training strategy: the first stage aligns vocal embeddings with the language model's input space, and the second stage trains the model to sequentially predict the appropriate response style and corresponding text.

The findings demonstrate substantial improvements across both automated metrics and human perception. Spoken-LLM outperformed text-only baselines and existing speech models in predicting response styles, achieving an emotion-prediction F1 score of 49.6 and a speaking-speed F1 score of 62.1 when utilizing chunk-level vocal embeddings. The framework achieved superior response diversity, recording a diversity score of 10.9 compared to 100.0 for text-only systems that generated identical text regardless of tone. In subjective listening tests, human evaluators overwhelmingly preferred Spoken-LLM over standard text-only systems, though a cascaded pipeline separating speech recognition and generation scored slightly higher in human naturalness due to subjective flexibility in conversational tone.

These results demonstrate that capturing vocal nuance is vital for deploying realistic conversational systems, directly reducing the risk of tone-deaf interactions in high-stakes customer-facing applications. The analysis also revealed that uncurated, synthetic dialogue data is insufficient on its own, as only about 33% of raw synthetic samples passed human naturalness screening. Employing a two-stage training approach that warms up on larger synthetic sets before fine-tuning on high-quality human-verified data proved essential to prevent overfitting.

Organizations developing voice-based conversational agents should adopt multimodal architectures that jointly process speech prosody and linguistic content rather than relying on text transcripts alone. Future technical roadmaps should focus on scaling human-verified datasets, incorporating spontaneous real-world conversational features such as laughter and turn-taking, and exploring direct speech-to-speech modeling that bypasses discrete style categories.

Key limitations include the modest size of the curated training set, which contained approximately 2,000 samples, and the reliance on synthetic speech rather than organic human recordings. Furthermore, while the current architecture shows robust performance, real-world deployment must account for acoustic noise and upstream transcription errors, both of which can degrade style classification and response quality.

Lin et al (2024).pdf

No sufficiently relevant recommendations were found.

Cover for Advancing Large Language Models to Capture Varied Speaking Styles and Respond Properly in Spoken Conversations

Abstract

In spoken dialogue, even if two current turns are the same sentence, their responses might still differ when they are spoken in different styles. The spoken styles, containing paralinguistic and prosodic information, mark the most significant difference between text and speech modality. When using text-only LLMs to model spoken dialogue, text-only LLMs cannot give different responses based on the speaking style of the current turn. In this paper, we focus on enabling LLMs to listen to the speaking styles and respond properly. Our goal is to teach the LLM that "even if the sentences are identical if they are spoken in different styles, their corresponding responses might be different". Since there is no suitable dataset for achieving this goal, we collect a speech-to-speech dataset, StyleTalk, with the following desired characteristics: when two current speeches have the same content but are spoken in different styles, their responses will be different. To teach LLMs to understand and respond properly to the speaking styles, we propose the Spoken-LLM framework that can model the linguistic content and the speaking styles. We train Spoken-LLM using the StyleTalk dataset and devise a two-stage training pipeline to help the Spoken-LLM better learn the speaking styles. Based on extensive experiments, we show that Spoken-LLM outperforms text-only baselines and prior speech LLMs methods.

Table of Contents

  • 1 Introduction
  • 2 Dataset: StyleTalk
  • 2.1 Overview
  • 2.2 Data Collection
  • 2.2.1 LLM for Data Generation
  • 2.2.2 Expressive Speech Synthesis
  • 2.2.3 Human Annotator Filtering
  • 2.3 Data split
  • 3 Spoken-LLM framework
  • 3.1 Overview
  • 3.2 Large Language Model
  • 3.3 Speech Style Encoder
  • 3.4 Spoken Dialogue Modeling
  • 3.5 Inference
  • 4 Experiments
  • 4.1 Baseline method
  • 4.2 Evaluation Metrics
  • 4.3 Main Results
  • 4.4 Subjective evaluation
  • 5 Analyses
  • 5.1 Same sentence in different speaking styles induce diverse responses
  • 5.2 Warmup pre-training and data quality
  • 5.3 Qualitative example
  • 6 Related works
  • 7 Conclusion
  • Limitation
  • Acknowledgement
  • References
  • Appendix
  • A Details of human annotators filtering
  • B Implementation details
  • C ASR prediction as input
  • D Style transition
  • E Diversity of current and responding styles
  • F Instruction
  • G Prompting GPT-4 for Data Generation
  • H Subjective evaluation
  • I Dataset license

Knowls

  1. Knowl 1 — StyleTalk pairs identical wording with style-dependent spoken responses

    experimental setup

    StyleTalk is a speech-to-speech dialogue dataset designed to test whether a model responds differently when the same current-turn sentence is spoken in different styles. Each dialogue set has a text dialogue context, a current sentence rendered with different speaking styles, and corresponding response speech. GPT-4 generated dialogue text and style annotations across 17 everyday topics; Microsoft Azure text-to-speech synthesized the speech using emotion, speed, and volume controls, with nine speakers (four male and five female). Human listeners heard the context, current speech, and three candidate responses, then selected the most natural response or indicated that none was natural. A sample was retained only when three listeners produced a majority choice and no listener selected the none-of-the-above option. Roughly 33% of generated samples passed this filtering.

  2. Knowl 2 — Spoken-LLM combines speech-style representations with a dialogue LLM

    model/method

    Spoken-LLM uses Llama 2-Chat 7B for dialogue generation and emotion2vec for representing the current speech’s paralinguistic and prosodic information. The Llama parameters are frozen while a LoRA adapter is trained; emotion2vec is also frozen. Speech features are mapped into the LLM input dimension by a trainable connector consisting of layer normalization and a linear projection. The connector accepts either a mean-pooled utterance representation or emotion2vec’s ten learned chunk embeddings. The text dialogue context and ground-truth transcript of the current turn are embedded as text, and the projected speech representation is combined with them as the LLM input. The model generates a response style—emotion, speed, and volume—followed by response text; an expressive text-to-speech system can then synthesize the response speech from those outputs.

  3. Knowl 3 — Two-stage training aligns speech features before response generation

    model/method

    Spoken-LLM is trained in two stages. First, the LLM is frozen and only the speech connector is trained to classify the current speech’s style attributes, aligning the speech representation with the LLM input space; this stage uses current-turn speech from the unfiltered generated data. Second, the connector and LoRA adapter are trained to generate the response style and then the response text, conditioned on the dialogue context and current turn. The second-stage objective is causal language-modeling cross-entropy over that ordered output. To make use of the limited human-filtered training data, the authors warm up on unfiltered generated samples and then fine-tune on the human-filtered set. The reported learning rates are 1e-3 for the first stage and 2e-4 for the second; training uses batch size 128 and LoRA rank 8.

  4. Knowl 4 — Spoken-LLM improves response-style prediction and is competitive on text metrics

    empirical result

    On the StyleTalk evaluation set using ground-truth current-turn transcripts, the authors compare text-only and cascaded text LLMs, ParalinGPT, and Spoken-LLM. Response text is scored with BLEU, ROUGE-L, METEOR, and BERT F1; response style is scored with weighted F1 for emotion, speed, and volume. The values below are reported scores. Spoken-LLM-chunk attains the highest emotion and speed F1 among the listed methods and matches the text-only upper-bound row on BLEU and BERT F1. Its volume F1 is lower than the text-only upper bound. Compared with Spoken-LLM-utt, the chunk variant is higher on BLEU, ROUGE-L, BERT F1, and all three style F1 scores, but lower on METEOR. The results also show that adding recognized current-style information to the text-only baseline improves response metrics.

    Method BLEU ROUGE-L METEOR BERT F1 F1 emotion F1 speed F1 volume
    Text-LLM (text-only) 3.1 16.2 17.4 75.3 17.5 37.1 41.9
    Text-LLM (cascaded) 3.2 17.3 19.1 76.0 37.5 52.9 65.6
    Text-LLM (upper bound) 4.0 17.9 19.6 76.3 40.2 53.5 65.8
    ParalinGPT-utt 3.1 16.8 18.5 75.9 32.3 51.9 64.8
    ParalinGPT-chunk 3.1 16.5 18.2 75.8 34.0 54.8 65.8
    Spoken-LLM-utt 2.8 16.6 20.2 75.8 47.4 61.5 56.5
    Spoken-LLM-chunk 4.0 17.8 19.4 76.3 49.6 62.1 61.1
  5. Knowl 5 — StyleTalk training and evaluation sets differ in size and style coverage

    data/table

    A sample is one current-speech and response-speech pair; a dialogue set groups samples sharing dialogue context and current-turn wording. The human-filtered training set contains 1,878 dialogue sets and 1,986 samples, while the evaluation set contains 486 dialogue sets and 981 samples. The additional unfiltered, LLM-generated set contains 5,777 dialogue sets and 16,472 samples and is not human-supervised. The dialogue-set counts below are grouped by the number of distinct current-turn speaking styles represented. Thus, although the evaluation set mostly has multiple styles for a given wording, most training dialogue sets have only one.

    Split One style Two styles Three styles
    Evaluation 16 445 25
    Training 1,770 108 0
    Unfiltered 0 859 4,918
  6. Knowl 6 — Warmup on unfiltered data followed by filtering improves overall metrics

    empirical result

    For Spoken-LLM-chunk, the authors compare training only on the human-filtered training set, only on unfiltered generated data, and warmup on unfiltered data followed by fine-tuning on the filtered set. The two-stage data sequence achieves the highest reported score on every listed metric. Training only on unfiltered data yields better response-text scores than training only on the smaller filtered set, but lower emotion, speed, and volume F1. The results support using the unfiltered data to learn task structure and language patterns before fine-tuning toward the human-filtered examples. Response-text metrics are BLEU, ROUGE-L, METEOR, and BERT F1; style metrics are weighted F1.

    Training data BLEU ROUGE-L METEOR BERT F1 F1 emotion F1 speed F1 volume
    Train 2.9 15.7 17.8 75.5 45.5 61.6 61.7
    Unfiltered 3.4 17.0 18.7 75.7 44.1 61.2 56.5
    Unfiltered then train 4.0 17.8 19.4 76.3 49.6 62.1 61.1
  7. Knowl 7 — Spoken-LLM generates more style-dependent response text than baselines

    empirical result

    The authors measure response-text diversity across evaluation dialogue sets using self-BLEU: for each set, they average the BLEU score between responses to different current speaking styles. Lower self-BLEU indicates more diverse responses. Spoken-LLM-chunk has a lower score than the cascaded text LLM and ParalinGPT-chunk, while the text-only LLM produces identical response content regardless of style and scores 100.0. The ground-truth responses score 8.2.

    Method Self-BLEU
    Ground truth 8.2
    Text-LLM (text-only) 100.0
    Text-LLM (cascaded) 11.2
    ParalinGPT-chunk 11.3
    Spoken-LLM-chunk 10.9
  8. Knowl 8 — Listeners prefer Spoken-LLM over text-only responses but slightly favor the cascade

    empirical result

    In a subjective comparison of generated response text and speech, three evaluators judged each of 200 samples in an A/B test. Listeners preferred Spoken-LLM-chunk over the text-only LLM by a large margin. They slightly preferred the cascaded text-LLM baseline over Spoken-LLM-chunk. The authors note that the two methods had similar response-text objective scores and that more than one response style can be appropriate for a given current turn, so a single reference style may not capture all natural responses.

  9. Knowl 9 — The StyleTalk and Spoken-LLM setup has limits in scale and speech realism

    limitation

    The authors identify several limits on the scope of the results. The human-filtered training set has only about 2,000 samples, creating risks of unstable training and overfitting. StyleTalk speech is synthesized with controlled Azure text-to-speech rather than collected as spontaneous speech, and emotion is represented with a single categorical label even though real speech may mix emotional attributes. Spoken-LLM predicts explicit style labels and response text for a separate expressive text-to-speech system rather than generating response speech directly. The turn-based setup also leaves out interaction behaviors such as backchannels, laughter, and turn-taking.

  10. Knowl 10 — Using ASR transcripts causes modest evaluation-score declines

    empirical result

    The authors test inference with Whisper base transcripts instead of ground-truth current-turn text. Whisper base has a 3.21% word error rate on current turns in the evaluation set. Both the cascaded text-LLM and Spoken-LLM-chunk show declines relative to their ground-truth-transcript results; the parenthesized values are the reported score changes. This is a supplementary input-condition test, not an analysis designed to resolve ASR error propagation.

    Method BLEU ROUGE-L METEOR BERT F1 F1 emotion F1 speed F1 volume
    Text-LLM (cascaded) 3.1 (-0.1) 16.9 (-0.4) 18.5 (-0.6) 75.9 (-0.1) 37.0 (-0.5) 52.5 (-0.4) 63.7 (-2.1)
    Spoken-LLM-chunk 3.3 (-0.7) 17.1 (-0.9) 19.0 (-0.4) 75.9 (-0.4) 47.5 (-2.1) 60.3 (-1.8) 57.4 (-3.7)

Coverage note — The supplementary emotion-transition visualizations and the detailed style-transition-pair diversity analysis are omitted because they are secondary diagnostics rather than central methods or headline evaluations.

References

  1. 1.Alexei Baevski, Arun Babu, Wei-Ning Hsu, and Michael Auli. 2023. Efficient self-supervised learning with contextualized target representations for vision, speech and language. In International Conference on Machine Learning, pages 1416–1429. PMLR.
  2. 2.Satanjeev Banerjee and Alon Lavie. 2005. Meteor: An automatic metric for mt evaluation with improved correlation with human judgments. In Proceedings of the acl workshop on intrinsic and extrinsic evaluation measures for machine translation and/or summarization, pages 65–72.
  3. 3.Zalán Borsos, Raphaël Marinier, Damien Vincent, Eugene Kharitonov, Olivier Pietquin, Matt Sharifi, Dominik Roblek, Olivier Teboul, David Grangier, Marco Tagliasacchi, et al. 2023. Audiolm: a language modeling approach to audio generation. IEEE/ACM Transactions on Audio, Speech, and Language Processing.
  4. 4.Carlos Busso, Murtaza Bulut, Chi-Chun Lee, Abe Kazemzadeh, Emily Mower, Samuel Kim, Jeannette N Chang, Sungbok Lee, and Shrikanth S Narayanan. 2008. Iemocap: Interactive emotional dyadic motion capture database. Language resources and evaluation, 42:335–359.
  5. 5.Carlos Busso, Srinivas Parthasarathy, Alec Burmania, Mohammed AbdelWahab, Najmeh Sadoughi, and Emily Mower Provost. 2016. Msp-improv: An acted corpus of dyadic interactions to study emotion perception. IEEE Transactions on Affective Computing, 8(1):67–80.
  6. 6.Santiago Castro, Devamanyu Hazarika, Verónica Pérez-Rosas, Roger Zimmermann, Rada Mihalcea, and Soujanya Poria. 2019. Towards multimodal sarcasm detection (an obviously perfect paper). In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4619–4629.
  7. 7.Eric Chen, Zhiyun Lu, Hao Xu, Liangliang Cao, Yu Zhang, and James Fan. 2020. A large scale speech sentiment corpus. In Proc. LREC, pages 6549–6555.
  8. 8.Ju-Chieh Chou, Chung-Ming Chien, Wei-Ning Hsu, Karen Livescu, Arun Babu, Alexis Conneau, Alexei Baevski, and Michael Auli. 2023. Toward joint language modeling for speech units and text. In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 6582–6593.
  9. 9.Yunfei Chu, Jin Xu, Xiaohuan Zhou, Qian Yang, Shiliang Zhang, Zhijie Yan, Chang Zhou, and Jingren Zhou. 2023. Qwen-audio: Advancing universal audio understanding via unified large-scale audio-language models. arXiv preprint arXiv:2311.07919.
  10. 10.Alexandre Défossez, Jade Copet, Gabriel Synnaeve, and Yossi Adi. 2022. High fidelity neural audio compression. arXiv preprint arXiv:2210.13438.
  11. 11.Soham Deshmukh, Benjamin Elizalde, Rita Singh, and Huaming Wang. 2023. Pengi: An audio language model for audio tasks. arXiv preprint arXiv:2305.11834.
  12. 12.Kevin Everson, Yile Gu, Huck Yang, Prashanth Gurunath Shivakumar, Guan-Ting Lin, Jari Kolehmainen, Ivan Bulyko, Ankur Gandhe, Shalini Ghosh, Wael Hamza, et al. 2024. Towards asr robust spoken language understanding through in-context learning with word confusion networks. arXiv preprint arXiv:2401.02921.
  13. 13.Mauajama Firdaus, Hardik Chauhan, Asif Ekbal, and Pushpak Bhattacharyya. 2020. Meisd: A multimodal multi-label emotion, intensity and sentiment dialogue dataset for emotion recognition and sentiment analysis in conversations. In Proceedings of the 28th international conference on computational linguistics, pages 4441–4453.
  14. 14.Yuan Gong, Alexander H Liu, Hongyin Luo, Leonid Karlinsky, and James Glass. 2023a. Joint audio and speech understanding. 2023 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU).
  15. 15.Yuan Gong, Hongyin Luo, Alexander H Liu, Leonid Karlinsky, and James Glass. 2023b. Listen, think, and understand. arXiv preprint arXiv:2305.10790.
  16. 16.Michael Hassid, Tal Remez, Tu Anh Nguyen, Itai Gat, Alexis Conneau, Felix Kreuk, Jade Copet, Alexandre Défossez, Gabriel Synnaeve, Emmanuel Dupoux, Roy Schwartz, and Yossi Adi. 2023. Textually pretrained speech language models. In Thirty-seventh Conference on Neural Information Processing Systems.
  17. 17.Mutian He and Philip N. Garner. 2023. Can ChatGPT Detect Intent? Evaluating Large Language Models for Spoken Language Understanding. In Proc. INTERSPEECH 2023, pages 1109–1113.
  18. 18.Edward J Hu, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, Weizhu Chen, et al. 2021. Lora: Low-rank adaptation of large language models. In International Conference on Learning Representations.
  19. 19.Rongjie Huang, Mingze Li, Dongchao Yang, Jiatong Shi, Xuankai Chang, Zhenhui Ye, Yuning Wu, Zhiqing Hong, Jiawei Huang, Jinglin Liu, et al. 2023. Audiogpt: Understanding and generating speech, music, sound, and talking head. arXiv preprint arXiv:2304.12995.
  20. 20.Eugene Kharitonov, Ann Lee, Adam Polyak, Yossi Adi, Jade Copet, Kushal Lakhotia, Tu-Anh Nguyen, Morgane Riviere, Abdelrahman Mohamed, Emmanuel Dupoux, et al. 2022. Text-free prosody-aware generative spoken language modeling. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 8666–8681.
  21. 21.Rithesh Kumar, Prem Seetharaman, Alejandro Luebs, Ishaan Kumar, and Kundan Kumar. 2023. High-fidelity audio compression with improved RVQGAN. In Thirty-seventh Conference on Neural Information Processing Systems.
  22. 22.Kushal Lakhotia, Evgeny Kharitonov, Wei-Ning Hsu, Yossi Adi, Adam Polyak, Benjamin Bolte, Tu-Anh Nguyen, Jade Copet, Alexei Baevski, Abdelrahman Mohamed, et al. 2021. On generative spoken language modeling from raw audio. Transactions of the Association for Computational Linguistics, 9:1336–1354.
  23. 23.Chin-Yew Lin. 2004. Rouge: A package for automatic evaluation of summaries. In Text summarization branches out, pages 74–81.
  24. 24.Guan-Ting Lin, Yung-Sung Chuang, Ho-Lam Chung, Shu wen Yang, Hsuan-Jui Chen, Shuyan Annie Dong, Shang-Wen Li, Abdelrahman Mohamed, Hung yi Lee, and Lin shan Lee. 2022. DUAL: Discrete Spoken Unit Adaptive Learning for Textless Spoken Question Answering. In Proc. Interspeech 2022, pages 5165–5169.
  25. 25.Guan-Ting Lin, Chi-Luen Feng, Wei-Ping Huang, Yuan Tseng, Tzu-Han Lin, Chen-An Li, Hung-yi Lee, and Nigel G Ward. 2023a. On the utility of self-supervised models for prosody-related tasks. In 2022 IEEE Spoken Language Technology Workshop (SLT), pages 1104–1111. IEEE.
  26. 26.Guan-Ting Lin, Prashanth Gurunath Shivakumar, Ankur Gandhe, Chao-Han Huck Yang, Yile Gu, Shalini Ghosh, Andreas Stolcke, Hung-yi Lee, and Ivan Bulyko. 2023b. Paralinguistics-enhanced large language modeling of spoken dialogue. arXiv preprint arXiv:2312.15316.
  27. 27.Ziyang Ma, Zhisheng Zheng, Jiaxin Ye, Jinchao Li, Zhifu Gao, Shiliang Zhang, and Xie Chen. 2023. emotion2vec: Self-supervised pre-training for speech emotion representation. arXiv preprint arXiv:2312.15185.
  28. 28.Soumi Maiti, Yifan Peng, Shukjae Choi, Jee-weon Jung, Xuankai Chang, and Shinji Watanabe. 2023. Voxtlm: unified decoder-only models for consolidating speech recognition/synthesis and speech/text continuation tasks. arXiv preprint arXiv:2309.07937.
  29. 29.Gary McKeown, Michel F Valstar, Roderick Cowie, and Maja Pantic. 2010. The semaine corpus of emotionally coloured character interactions. In 2010 IEEE International Conference on Multimedia and Expo, pages 1079–1084. IEEE.
  30. 30.Kentaro Mitsui, Yukiya Hono, and Kei Sawada. 2023. Towards human-like spoken dialogue generation between ai agents from written dialogue. arXiv preprint arXiv:2310.01088.
  31. 31.Abdelrahman Mohamed, Hung-yi Lee, Lasse Borgholt, Jakob D. Havtorn, Joakim Edin, Christian Igel, Katrin Kirchhoff, Shang-Wen Li, Karen Livescu, Lars Maaløe, Tara N. Sainath, and Shinji Watanabe. 2022. Self-supervised speech representation learning: A review. IEEE Journal of Selected Topics in Signal Processing, 16(6):1179–1210.
  32. 32.Eliya Nachmani, Alon Levkovitch, Julian Salazar, Chulayutsh Asawaroengchai, Soroosh Mariooryad, RJ Skerry-Ryan, and Michelle Tadmor Ramanovich. 2023. Lms with a voice: Spoken language modeling beyond speech tokens. arXiv preprint arXiv:2305.15255.
  33. 33.Tu Anh Nguyen, Eugene Kharitonov, Jade Copet, Yossi Adi, Wei-Ning Hsu, Ali Elkahky, Paden Tomasello, Robin Algayres, Benoit Sagot, Abdelrahman Mohamed, et al. 2023. Generative spoken dialogue language modeling. Transactions of the Association for Computational Linguistics, 11:250–266.
  34. 34.OpenAI. 2023. Gpt-4 technical report.
  35. 35.Jing Pan, Jian Wu, Yashesh Gaur, Sunit Sivasankaran, Zhuo Chen, Shujie Liu, and Jinyu Li. 2023. Cosmic: Data efficient instruction-tuning for speech in-context learning. arXiv preprint arXiv:2311.02248.
  36. 36.Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th annual meeting of the Association for Computational Linguistics, pages 311–318.
  37. 37.Soujanya Poria, Devamanyu Hazarika, Navonil Majumder, Gautam Naik, Erik Cambria, and Rada Mihalcea. 2019. Meld: A multimodal multi-party dataset for emotion recognition in conversations. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 527–536.
  38. 38.Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine Mcleavey, and Ilya Sutskever. 2023. Robust speech recognition via large-scale weak supervision. In Proceedings of the 40th International Conference on Machine Learning, volume 202 of Proceedings of Machine Learning Research, pages 28492–28518. PMLR.
  39. 39.Paul K Rubenstein, Chulayuth Asawaroengchai, Duc Dung Nguyen, Ankur Bapna, Zalán Borsos, Félix de Chaumont Quitry, Peter Chen, Dalia El Badawy, Wei Han, Eugene Kharitonov, et al. 2023. Audiopalm: A large language model that can speak and listen. arXiv preprint arXiv:2306.12925.
  40. 40.Yongliang Shen, Kaitao Song, Xu Tan, Dongsheng Li, Weiming Lu, and Yueting Zhuang. 2023. Hugginggpt: Solving ai tasks with chatgpt and its friends in huggingface. arXiv preprint arXiv:2303.17580.
  41. 41.Changli Tang, Wenyi Yu, Guangzhi Sun, Xianzhao Chen, Tian Tan, Wei Li, Lu Lu, Zejun Ma, and Chao Zhang. 2023. Salmonn: Towards generic hearing abilities for large language models. arXiv preprint arXiv:2310.13289.
  42. 42.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288.
  43. 43.Hsiang-Sheng Tsai, Heng-Jui Chang, Wen-Chin Huang, Zili Huang, Kushal Lakhotia, Shu-wen Yang, Shuyan Dong, Andy Liu, Cheng-I Lai, Jiatong Shi, et al. 2022. Superb-sg: Enhanced speech processing universal performance benchmark for semantic and generative capabilities. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 8479–8492.
  44. 44.Tianrui Wang, Long Zhou, Ziqiang Zhang, Yu Wu, Shujie Liu, Yashesh Gaur, Zhuo Chen, Jinyu Li, and Furu Wei. 2023. Viola: Unified codec language models for speech recognition, synthesis, and translation. arXiv preprint arXiv:2305.16107.
  45. 45.Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, et al. 2022. Emergent abilities of large language models. arXiv preprint arXiv:2206.07682.
  46. 46.Jian Wu, Yashesh Gaur, Zhuo Chen, Long Zhou, Yimeng Zhu, Tianrui Wang, Jinyu Li, Shujie Liu, Bo Ren, Linquan Liu, et al. 2023a. On decoder-only architecture for speech-to-text and large language model integration. In 2023 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), pages 1–8. IEEE.
  47. 47.Yi-Chiao Wu, Israel D Gebru, Dejan Markovic, and ´ Alexander Richard. 2023b. Audiodec: An open-source streaming high-fidelity neural audio codec. In ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 1–5. IEEE.
  48. 48.Hongfei Xue, Yuhao Liang, Bingshen Mu, Shiliang Zhang, Qian Chen, and Lei Xie. 2023. E-chat: Emotion-sensitive spoken dialogue system with large language models. arXiv preprint arXiv:2401.00475.
  49. 49.Dongchao Yang, Songxiang Liu, Rongjie Huang, Jinchuan Tian, Chao Weng, and Yuexian Zou. 2023. Hifi-codec: Group-residual vector quantization for high fidelity audio codec. arXiv preprint arXiv:2305.02765.
  50. 50.Shu-Wen Yang, Po-Han Chi, Yung-Sung Chuang, Cheng-I Jeff Lai, Kushal Lakhotia, Yist Y. Lin, Andy T. Liu, Jiatong Shi, Xuankai Chang, Guan-Ting Lin, Tzu-Hsien Huang, Wei-Cheng Tseng, Ko tik Lee, Da-Rong Liu, Zili Huang, Shuyan Dong, Shang-Wen Li, Shinji Watanabe, Abdelrahman Mohamed, and Hung yi Lee. 2021. SUPERB: Speech Processing Universal PERformance Benchmark. In Proc. Interspeech 2021, pages 1194–1198.
  51. 51.Chenyu You, Nuo Chen, Fenglin Liu, Shen Ge, Xian Wu, and Yuexian Zou. 2022. End-to-end spoken conversational question answering: Task, dataset and model. In Findings of the Association for Computational Linguistics: NAACL 2022, pages 1219–1232.
  52. 52.Neil Zeghidour, Alejandro Luebs, Ahmed Omran, Jan Skoglund, and Marco Tagliasacchi. 2021. Soundstream: An end-to-end neural audio codec. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 30:495–507.
  53. 53.Dong Zhang, Shimin Li, Xin Zhang, Jun Zhan, Pengyu Wang, Yaqian Zhou, and Xipeng Qiu. 2023. SpeechGPT: Empowering large language models with intrinsic cross-modal conversational abilities. In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 15757–15773, Singapore. Association for Computational Linguistics.
  54. 54.Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi. 2019. Bertscore: Evaluating text generation with bert. In International Conference on Learning Representations.
  55. 55.Yaoming Zhu, Sidi Lu, Lei Zheng, Jiaxian Guo, Weinan Zhang, Jun Wang, and Yong Yu. 2018. Texygen: A benchmarking platform for text generation models. In The 41st international ACM SIGIR conference on research & development in information retrieval, pages 1097–1100.

Citation

MLA
Lin, G.-T., et al. “Advancing Large Language Models to Capture Varied Speaking Styles and Respond Properly in Spoken Conversations”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 6626–42, https://doi.org/10.18653/v1/2024.acl-long.358.
APA
Lin, G.-T., Chiang, C.-H., & Lee, H.-. yi . (2024). Advancing Large Language Models to Capture Varied Speaking Styles and Respond Properly in Spoken Conversations. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 6626–6642. https://doi.org/10.18653/v1/2024.acl-long.358
Chicago
Lin, G.-T., C.-H. Chiang, and H.-. yi . Lee. 2024. “Advancing Large Language Models to Capture Varied Speaking Styles and Respond Properly in Spoken Conversations”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 6626–42. https://doi.org/10.18653/v1/2024.acl-long.358.
Harvard
Lin, G.-T., Chiang, C.-H. and Lee, H.-. yi . (2024) “Advancing Large Language Models to Capture Varied Speaking Styles and Respond Properly in Spoken Conversations”, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 6626–6642. Available at: https://doi.org/10.18653/v1/2024.acl-long.358.
Vancouver
1. Lin G-T, Chiang C-H, Lee H-yi (2024) Advancing Large Language Models to Capture Varied Speaking Styles and Respond Properly in Spoken Conversations. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 6626–6642

BibTeX

@inproceedings{lin-etal-2024-advancing,
    title = "Advancing Large Language Models to Capture Varied Speaking Styles and Respond Properly in Spoken Conversations",
    author = "Lin, Guan-Ting  and
      Chiang, Cheng-Han  and
      Lee, Hung-yi",
    editor = "Ku, Lun-Wei  and
      Martins, Andre  and
      Srikumar, Vivek",
    booktitle = "Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = aug,
    year = "2024",
    address = "Bangkok, Thailand",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.acl-long.358/",
    doi = "10.18653/v1/2024.acl-long.358",
    pages = "6626--6642"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/