SD-Eval: A Benchmark Dataset for Spoken Dialogue Understanding Beyond Words

Junyi AoYuancheng WangXiaohai TianDekun ChenJun ZhangLu LuYuxuan WangHaizhou LiZhizheng Wu

article2024NeurIPS90 citations

Introduces SD-Eval, an open-source benchmark and training resource designed to evaluate how effectively speech language models integrate speaker emotion, accent, age, and background sounds into dialogue generation.

Listen

Spoken communication conveys rich contextual cues beyond literal words, including vocal emotions, accents, speaker age, and environmental background noises. Although modern speech-enabled artificial intelligence systems can transcribe words and perceive audio features, they frequently fail to generate context-aware, empathetic responses in real conversations. This limitation stems from a lack of standard task definitions, open-source datasets, and specialized evaluation frameworks designed to assess conversational responses conditioned on non-verbal audio cues.

The article introduces SD-Eval, a benchmark dataset designed to evaluate spoken dialogue understanding and response generation across multiple dimensions beyond text alone. The objective is to establish standardized resources and metrics that promote the development of intelligent dialogue systems capable of adapting responses to vocal and environmental nuances.

The researchers constructed the SD-Eval benchmark using 7,303 speech utterances totaling 8.76 hours across four evaluation sub-tasks: emotion, accent, age, and environmental sounds. The data was curated from eight public speech datasets, incorporating both real recordings and targeted synthetic speech. To support baseline model development, the researchers assembled a corresponding training dataset containing 724,400 utterances totaling 1,052.72 hours across eleven public datasets. They evaluated baseline systems—including text-cascaded models, an end-to-end speech large language model, and several open-source audio language models—using traditional text-matching metrics, advanced large language model judges, and human evaluations across 200 sampled conversations.

The findings demonstrate that the end-to-end speech model, which ingests raw audio representations directly, consistently outperformed traditional cascaded systems that rely solely on automated text transcription. For example, on the emotion subset, the end-to-end model achieved a response quality score of 5.30 compared to 4.47 for the cascaded baseline when rated by an advanced model judge, reflecting its ability to pick up non-verbal cues directly from speech. Providing explicit, high-quality ground-truth contextual labels produced the highest overall performance, demonstrating that input quality significantly influences conversational appropriateness. Furthermore, off-the-shelf open-source models performed poorly on these nuanced conversational tasks unless specifically operated in dedicated voice-chat modes. Methodologically, automated evaluations conducted using advanced large language models aligned substantially better with human judgments than traditional metrics like BLEU or ROUGE, with the top model judge achieving a strong overall correlation of 0.613 compared to 0.208 for BLEU-4.

These results indicate that conversational systems relying solely on speech-to-text pipelines risk misinterpreting user intent by ignoring paralinguistic and ambient context. Transitioning toward end-to-end audio processing or incorporating explicit acoustic feature recognition is essential for deploying conversational artificial intelligence in customer service, healthcare, and voice assistance where empathy and situational awareness are critical. Additionally, organizations can adopt large language model judges as reliable, scalable evaluation proxies, substantially reducing the high financial and operational costs associated with manual human quality reviews.

Stakeholders developing voice-based interactive systems should transition from pure text-cascaded architectures to end-to-end multi-modal models trained explicitly on diverse acoustic contexts. Development teams should expand training data to include broader demographic characteristics, explore feature-disentanglement methods, and adopt model-based evaluation frameworks to monitor response quality. Readers should note that SD-Eval is currently constrained to single-turn, speech-to-text dialogues and does not evaluate speaker gender or multi-turn speech-to-speech interactions, representing key areas for future research.

Cover for SD-Eval: A Benchmark Dataset for Spoken Dialogue Understanding Beyond Words

Abstract

Speech encompasses a wealth of information, including but not limited to content, paralinguistic, and environmental information. This comprehensive nature of speech significantly impacts communication and is crucial for human-computer interaction. Chat-Oriented Large Language Models (LLMs), known for their general-purpose assistance capabilities, have evolved to handle multi-modal inputs, including speech. Although these models can be adept at recognizing and analyzing speech, they often fall short of generating appropriate responses. We argue that this is due to the lack of principles on task definition and model development, which requires open-source datasets and metrics suitable for model evaluation. To bridge the gap, we present SD-Eval, a benchmark dataset aimed at multidimensional evaluation of spoken dialogue understanding and generation. SD-Eval focuses on paralinguistic and environmental information and includes 7,303 utterances, amounting to 8.76 hours of speech data. The data is aggregated from eight public datasets, representing four perspectives: emotion, accent, age, and background sound. To assess the SD-Eval benchmark dataset, we implement three different models and construct a training set following a process similar to that of SD-Eval. The training set contains 1,052.72 hours of speech data and 724.4k utterances. We also conduct a comprehensive evaluation using objective evaluation methods (e.g. BLEU and ROUGE), subjective evaluations, and LLM-based metrics for the generated responses. Models conditioned with paralinguistic and environmental information outperform their counterparts in both objective and subjective measures. Moreover, experiments demonstrate that LLM-based metrics show a higher correlation with human evaluation compared to traditional metrics. We open-source SD-Eval at https://github.com/amphionspace/SD-Eval.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 SD-Eval Benchmark Dataset
  • 3.1 Dataset Construction
  • 3.2 Dataset Statistics
  • 4 Benchmark Experiments
  • 4.1 Training Set
  • 4.2 Models
  • 4.3 Evaluation Metrics
  • 4.4 Experimental Setup
  • 4.5 Main Results
  • 4.6 Analysis
  • 5 Conclusion
  • 6 Limitations and Future Work
  • Acknowledgments
  • References
  • A Appendix
  • A.1 Statistics of Training Set
  • A.2 Zero-shot TTS Model
  • A.3 Prompts for Generating Responses
  • A.3.1 Prompts for Training Set
  • A.3.2 Prompts for SD-Eval
  • A.4 Prompts for LLM Evaluation
  • A.5 Prompts for Generating Dialogue Data of test-env
  • A.6 Text Instruction Prompts for Open-Sourced Models

Knowls

  1. Knowl 1 — SD-Eval benchmark scope and task definition

    definition

    SD-Eval is a benchmark for evaluating spoken-dialogue systems that must generate appropriate text responses from speech containing information beyond the literal words. Its initial release evaluates single-turn speech-to-text dialogue across four independent subsets: test-emo for emotion, test-acc for accent, test-age for speaker age, and test-env for background environment sounds. The intended longer-term target is multidimensional speech-to-speech conversation, but the released benchmark evaluates generated responses at the text level.

  2. Knowl 2 — Construction and annotation pipeline for SD-Eval

    model/method

    SD-Eval combines real and synthesized speech from eight public datasets. Emotion data comes from RAVDESS, MEAD, and JL Corpus, which provide similar content spoken with different emotions. Accent data comes from VCTK and Common Voice. Age data uses children’s conversational text from MyST and adult speech synthesized with an internal zero-shot TTS model trained on Libri-light; a randomly selected LibriSpeech test-clean utterance supplies the synthesis prompt for each text. Environment data combines LibriSpeech test-clean speech with randomly selected AudioCaps recordings from seven scene categories and also includes GPT-4-Turbo-generated dialogue text rendered by the TTS model.

    Labels are normalized to nine accents—England, Scottish, Irish, Welsh, Northern Irish, American, Canadian, Australian, and New Zealand—five emotions—sad, happy, angry, disgust, and fear—and two age groups, child and adult. Neutral and surprise emotion samples are removed because neutral responses can depend mainly on content and surprise can correspond to different sentiments. The seven environment categories are driving, children’s voice, sea beach, raining or thundering, bells, sports center, and bus or subway.

    The filtering process first uses a GPT-4-Turbo prompt to remove ambiguous or insufficiently contextual utterances, followed by evaluation by three human annotators. Human annotators additionally remove environment examples whose background sound is incorrect. Emotion examples are also filtered when the transcript sentiment and the speech emotion have the same positive or negative polarity, so that the effect of vocal emotion is more salient. A punctuation-restoration model is applied to transcripts lacking punctuation. GPT-4o then generates five diverse candidate responses for every benchmark utterance, conditioned on the utterance content and its emotion, accent, age, or environment label.

  3. Knowl 3 — SD-Eval dataset composition

    data/table

    SD-Eval contains 7,303 utterances totaling 8.76 hours. The four subsets vary in size and target information as follows:

    Type Hours Utterances Sources Labels
    Emotion (`test-emo`) 1.11 1,289 RAVDESS, MEAD, JL Corpus Sad, Angry, Fear, Disgust, Happy
    Accent (`test-acc`) 5.34 4,310 VCTK, Common Voice England, Scottish, Northern Irish, Welsh, Irish, American, Canadian, Australian, New Zealand
    Environment (`test-env`) 0.74 690 LibriSpeech, AudioCaps, synthesized speech Driving, Children’s Voice, Sea Beach, Raining or Thundering, Bells, Sports Center, Bus or Subway
    Age (`test-age`) 1.57 1,014 MyST, synthesized speech Adult, Child
    Summary 8.76 7,303 – –

    The composition is intended to test whether a response changes appropriately with four kinds of speech-carried information rather than only with the transcript content.

  4. Knowl 4 — Training corpus used to assess the benchmark

    data/table

    The authors construct a separate training corpus from eleven open datasets using a process similar to SD-Eval. Unlike the benchmark construction, the training pipeline removes inadequate or ambiguous labels without the full filtering procedure and generates one response per utterance, except that environment examples receive five responses for augmentation. The appendix reports the following statistics:

    Type Hours Utterances Sources Labels Generator
    Emotion 120.60 100.5k MSP-Podcast, IEMOCAP, MELD, EmoV-DB, ESD, CREMA-D Angry, Contempt, Disgust, Fear, Happy, Neutral, Sad, Surprise, Frustrated, Excited, Amused, Sleepiness GPT-3.5-Turbo
    Accent 759.75 508.6k UK-Ireland dataset, VCTK, Common Voice England, Scottish, Northern Irish, Welsh, Irish, American, Canadian, Australian, New Zealand GPT-4o
    Environment 32.06 47.1k LibriSpeech, AudioCaps, synthesized speech Driving, Children’s Voice, Sea Beach, Raining or Thundering, Bells, Sports Center, Shopping Center, Bus or Subway GPT-4-Turbo
    Age 140.31 73.2k MyST Child GPT-3.5-Turbo
    Summary 1,052.72 729.4k – – –

    The abstract and main text describe the training set as containing 724.4k utterances, whereas the appendix statistics report 729.4k utterances.

  5. Knowl 5 — Speech-dialogue models evaluated by the benchmark

    model/method

    The benchmark compares three internally trained systems with three off-the-shelf speech-language systems. The Cascade LLM first applies Whisper large-v3 automatic speech recognition and then gives the resulting transcript to a frozen 7-billion-parameter InternLM2-chat model adapted with trainable LoRA parameters; it therefore has access primarily to recognized linguistic content. The Vanilla Speech LLM (VS-LLM) feeds Whisper large-v3 speech representations directly through a trainable adaptor into the same frozen InternLM2-chat model with LoRA. The adaptor has two linear layers, GELU after the first layer, and two-dimensional average pooling after the second layer to down-sample the speech representation.

    The LLM Upper Bound uses the ground-truth transcript concatenated with the ground-truth paralinguistic or environmental label, such as How are you?<Emotion:Happy>. It measures performance when content and target non-lexical information are supplied without speech-recognition or perception errors. The external systems are Qwen-Audio, Qwen2-Audio in audio-analysis mode, Qwen2-Audio in voice-chat mode, and SALMONN; each receives a text instruction prompting it to generate a response from the audio.

    Internally trained systems use AdamW with learning rate 2×10−42\times10^{-4}, two epochs, batch size 16, and 16 A100 GPUs. The InternLM2 LoRA adaptor uses rank 512 and α=256\alpha=256; the Whisper encoder adaptor uses rank 64 and α=16\alpha=16.

  6. Knowl 6 — Objective and human evaluation protocol

    experimental setup

    Generated responses are scored with reference-based lexical metrics—BLEU-4, ROUGE-L, and METEOR—and embedding-based BERTScore using a RoBERTa-large model. The authors also introduce a reference-free LLM-judge metric: a judge assigns a score from 1 to 10 to an individual response after considering naturalness, coherence, engagingness, groundedness, and whether the response appropriately uses the relevant emotion, accent, age, or background sound. Judges include GPT-4o, Yi-1.5-34B-Chat, Qwen2-57B-A14B-Instruct, and Gemma-2-27B-it; quantized Q6_K inference through llama.cpp allows evaluation on CPUs or a single A100 GPU.

    Human evaluation uses 200 randomly selected utterances, with 50 from each of the four subsets. Each utterance is paired with responses from Cascade LLM, VS-LLM, and the LLM Upper Bound. Every valid sample is rated by at least three annotators, yielding at least 120 valid evaluated samples per subset.

  7. Knowl 7 — Main benchmark results across the four information types

    data/table

    The following results compare the four external systems with the three internally trained systems. BLEU-4, ROUGE-L, and METEOR are lexical scores; BERTScore is reported on a 0–100-style scale; each LLM-judge score is on a 1–10 scale; human scores use the same 1–10 scale. Human scores were collected only for the three internally trained systems.

    Subset Model BLEU-4 ROUGE-L METEOR BERTScore Yi-1.5 Qwen2 Gemma GPT-4o Human
    test-emo SALMONN 2.48 16.57 18.97 86.20 4.98 3.35 2.32 2.61 –
    Qwen-Audio 3.93 19.02 16.82 86.59 4.19 2.35 2.02 2.24 –
    Qwen2-Audio-AA 3.01 16.82 17.51 86.17 4.75 2.52 2.21 2.33 –
    Qwen2-Audio-VC 2.21 14.57 22.08 85.41 5.88 3.83 2.93 3.25 –
    Cascade LLM 4.66 21.98 21.70 87.93 5.67 3.86 2.35 4.47 5.05
    VS-LLM 8.29 25.52 27.23 89.48 6.40 4.56 4.03 5.30 6.31
    LLM Upper Bound 12.35 26.08 28.27 89.77 7.03 5.82 6.46 6.74 7.29
    test-acc SALMONN 7.50 22.22 21.23 87.53 5.27 6.16 3.16 2.93 –
    Qwen-Audio 4.52 17.15 17.78 85.59 3.48 3.45 1.86 1.72 –
    Qwen2-Audio-AA 7.26 21.80 19.68 87.68 5.04 6.13 3.01 2.54 –
    Qwen2-Audio-VC 3.47 17.46 23.77 86.26 5.96 6.20 3.94 4.37 –
    Cascade LLM 14.51 30.53 34.13 89.66 7.23 7.32 5.65 6.62 6.71
    VS-LLM 17.98 33.06 37.65 90.08 7.82 7.65 6.59 7.85 7.95
    LLM Upper Bound 18.35 33.48 38.27 90.23 7.85 7.75 6.73 8.02 8.30
    test-age SALMONN 10.03 24.95 23.55 88.10 5.41 4.66 3.14 3.35 –
    Qwen-Audio 7.28 23.09 21.80 86.72 4.43 3.98 2.25 2.50 –
    Qwen2-Audio-AA 6.81 22.72 20.51 87.47 5.19 4.58 3.01 3.14 –
    Qwen2-Audio-VC 5.64 18.90 28.23 86.70 7.03 5.92 4.48 5.06 –
    Cascade LLM 15.36 31.96 31.99 90.08 7.22 7.16 6.46 4.47 6.51
    VS-LLM 17.22 34.17 33.78 90.63 7.74 7.39 7.25 7.95 7.11
    LLM Upper Bound 18.78 35.62 36.01 91.00 7.82 7.54 7.40 8.25 7.44
    test-env SALMONN 2.87 16.53 21.37 86.71 4.70 5.00 3.40 3.56 –
    Qwen-Audio 2.37 16.83 17.50 85.81 3.77 1.86 2.16 2.14 –
    Qwen2-Audio-AA 2.97 16.32 19.84 86.50 4.52 5.02 3.49 3.50 –
    Qwen2-Audio-VC 2.06 12.35 23.40 85.17 6.30 6.21 4.85 5.30 –
    Cascade LLM 5.44 21.75 26.41 88.22 6.03 5.84 5.31 5.66 6.62
    VS-LLM 9.42 25.85 28.27 89.23 6.14 5.88 5.10 5.82 7.11
    LLM Upper Bound 11.72 27.95 31.50 89.73 7.14 7.14 6.25 7.40 8.13

    VS-LLM outperforms Cascade LLM on every reported metric in all four subsets, showing the benefit of directly processing speech rather than relying only on ASR text. The LLM Upper Bound remains better than VS-LLM, indicating that ground-truth transcripts and labels are still more effective than implicitly recovering the same information from speech. Among the external systems, Qwen2-Audio voice-chat mode is generally strongest, but remains below the internally trained VS-LLM.

  8. Knowl 8 — Input-quality ablation on emotional dialogue

    data/table

    An ablation on test-emo separates the effects of speech input, transcript quality, and emotion-label quality. Speech LLM denotes a model receiving speech; LLM denotes a text model. ASR and GT indicate an automatic or ground-truth transcript, respectively. SER and GT under Emotion Label indicate an emotion label from emotion2vec or the ground-truth emotion, respectively.

    Index Model Type Transcript Emotion Label BLEU-4 ROUGE-L METEOR BERTScore Yi-1.5 Qwen2 Gemma GPT-4o
    1 Speech LLM N/A N/A 8.29 25.52 27.23 89.48 6.40 4.56 4.03 5.30
    2 Speech LLM N/A SER 10.37 27.29 28.59 89.81 6.76 5.26 4.37 6.11
    3 Speech LLM N/A GT 10.21 27.22 28.45 89.85 6.96 5.45 4.45 6.41
    4 LLM ASR N/A 4.66 21.98 21.70 87.93 5.67 3.86 2.35 4.47
    5 LLM ASR SER 11.37 26.03 27.66 89.66 6.74 5.35 5.62 6.13
    6 LLM GT SER 11.85 26.05 27.78 89.75 6.85 5.53 6.07 6.38
    7 LLM ASR GT 11.85 26.03 28.19 89.68 6.96 5.64 6.04 6.47
    8 LLM GT GT 12.35 26.08 28.27 89.77 7.03 5.82 6.46 6.74

    For text-based systems, replacing ASR transcripts with ground-truth transcripts improves every metric: the fully ground-truth system reaches BLEU-4 12.35 and GPT-4o 6.74, compared with 4.66 and 4.47 when no emotion label is supplied. Replacing SER emotion labels with ground-truth labels also improves all metrics. For the speech-only system, adding an automatically predicted emotion label raises BLEU-4 from 8.29 to 10.37, and using the ground-truth emotion label produces the strongest speech-input result on most metrics, including GPT-4o 6.41.

  9. Knowl 9 — Correlation of evaluation metrics with human judgments

    data/table

    The authors compute dataset-level Spearman correlation ρ\rho and Kendall-Tau correlation τ\tau between each automatic metric and human ratings. The LLM judges correlate more strongly with human judgments than BLEU-4, ROUGE-L, METEOR, or BERTScore in the overall evaluation and across the four subsets. GPT-4o is the strongest overall judge, with overall ρ=0.613\rho=0.613 and τ=0.468\tau=0.468.

    Metric test-emo test-acc test-age test-env Overall
    ρ\rho τ\tau ρ\rho τ\tau ρ\rho τ\tau ρ\rho τ\tau ρ\rho τ\tau
    BLEU-4 0.211 0.170 0.197 0.156 0.316 0.232 0.169 0.136 0.208 0.160
    ROUGE-L 0.203 0.140 0.263 0.192 0.318 0.215 0.216 0.144 0.258 0.176
    METEOR 0.249 0.168 0.272 0.194 0.350 0.242 0.315 0.214 0.336 0.229
    BERTScore 0.378 0.269 0.252 0.184 0.321 0.216 0.286 0.199 0.291 0.199
    Yi-1.5 0.641 0.493 0.435 0.345 0.356 0.289 0.558 0.441 0.492 0.381
    Qwen2 0.618 0.456 0.198 0.161 0.347 0.278 0.449 0.352 0.474 0.362
    Gemma 0.639 0.493 0.375 0.304 0.448 0.356 0.563 0.439 0.492 0.380
    GPT-4o 0.731 0.568 0.659 0.541 0.474 0.356 0.577 0.459 0.613 0.468

    For comparison, the strongest traditional overall correlation is METEOR with ρ=0.336\rho=0.336 and τ=0.229\tau=0.229, while GPT-4o reaches ρ=0.613\rho=0.613 and τ=0.468\tau=0.468. This supports using LLM-based, criterion-driven judges for the open-ended response task, although the correlations vary by information type.

  10. Knowl 10 — Stated limitations and future extensions

    limitation

    The released SD-Eval benchmark evaluates only speech-to-text dialogues, so it cannot assess whether a system produces an appropriate spoken response. It is also restricted to single-turn interactions and therefore does not test context accumulation or response quality in multi-turn conversations. Finally, its four task dimensions—emotion, accent, age, and environmental sound—do not include other speaker attributes such as gender. The authors identify multi-turn speech-to-speech evaluation and broader coverage of speech information as future directions.

Coverage note — The supplementary zero-shot TTS comparison and the full prompt listings were omitted because they support synthetic-data generation and evaluation implementation but are not independent benchmark contributions.

References

  1. 1.Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. Gpt-4 technical report. arXiv preprint arXiv:2303.08774, 2023.
  2. 2.Adaeze Adigwe, Noé Tits, Kevin El Haddad, Sarah Ostadabbas, and Thierry Dutoit. The emotional voices database: Towards controlling the emotion dimension in voice generation systems. arXiv preprint arXiv:1806.09514, 2018.
  3. 3.R. Ardila, M. Branson, K. Davis, M. Henretty, M. Kohler, J. Meyer, R. Morais, L. Saunders, F. M. Tyers, and G. Weber. Common voice: A massively-multilingual speech corpus. In Proceedings of the 12th Conference on Language Resources and Evaluation (LREC 2020), pages 4211–4215, 2020.
  4. 4.Satanjeev Banerjee and Alon Lavie. METEOR: An automatic metric for MT evaluation with improved correlation with human judgments. In Jade Goldstein, Alon Lavie, Chin-Yew Lin, and Clare Voss, editors, Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and/or Summarization, pages 65–72, Ann Arbor, Michigan, June 2005. Association for Computational Linguistics. URL https://aclanthology.org/W05-0909.
  5. 5.Carlos Busso, Murtaza Bulut, Chi-Chun Lee, Abe Kazemzadeh, Emily Mower, Samuel Kim, Jeannette N Chang, Sungbok Lee, and Shrikanth S Narayanan. Iemocap: Interactive emotional dyadic motion capture database. Language resources and evaluation, 42:335–359, 2008.
  6. 6.Zheng Cai, Maosong Cao, Haojiong Chen, Kai Chen, Keyu Chen, Xin Chen, Xun Chen, Zehui Chen, Zhi Chen, Pei Chu, et al. Internlm2 technical report. arXiv preprint arXiv:2403.17297, 2024.
  7. 7.Houwei Cao, David G. Cooper, Michael K. Keutmann, Ruben C. Gur, Ani Nenkova, and Ragini Verma. Crema-d: Crowd-sourced emotional multimodal actors dataset. IEEE Transactions on Affective Computing, 5(4):377–390, 2014. doi: 10.1109/TAFFC.2014.2336244.
  8. 8.Sanyuan Chen, Chengyi Wang, Zhengyang Chen, Yu Wu, Shujie Liu, Zhuo Chen, Jinyu Li, Naoyuki Kanda, Takuya Yoshioka, Xiong Xiao, et al. Wavlm: Large-scale self-supervised pre-training for full stack speech processing. IEEE Journal of Selected Topics in Signal Processing, 16(6):1505–1518, 2022.
  9. 9.Sanyuan Chen, Yu Wu, Chengyi Wang, Shujie Liu, Daniel Tompkins, Zhuo Chen, and Furu Wei. Beats: Audio pre-training with acoustic tokenizers. arXiv preprint arXiv:2212.09058, 2022.
  10. 10.Yunfei Chu, Jin Xu, Xiaohuan Zhou, Qian Yang, Shiliang Zhang, Zhijie Yan, Chang Zhou, and Jingren Zhou. Qwen-audio: Advancing universal audio understanding via unified large-scale audio-language models. arXiv preprint arXiv:2311.07919, 2023.
  11. 11.Yunfei Chu, Jin Xu, Qian Yang, Haojie Wei, Xipin Wei, Zhifang Guo, Yichong Leng, Yuanjun Lv, Jinzheng He, Junyang Lin, et al. Qwen2-audio technical report. arXiv preprint arXiv:2407.10759, 2024.
  12. 12.P R Cohen and S L Oviatt. The role of voice input for human-machine communication. Proceedings of the National Academy of Sciences, 92(22):9921–9927, 1995. doi: 10.1073/pnas.92.22.9921. URL https://www.pnas.org/doi/abs/10.1073/pnas.92.22.9921.
  13. 13.XTuner Contributors. Xtuner: A toolkit for efficiently fine-tuning llm. https://github.com/InternLM/xtuner, 2023.
  14. 14.R. Cowie, E. Douglas-Cowie, N. Tsapatsoulis, G. Votsis, S. Kollias, W. Fellenz, and J.G. Taylor. Emotion recognition in human-computer interaction. IEEE Signal Processing Magazine, 18(1):32–80, 2001. doi: 10.1109/79.911197.
  15. 15.Isin Demirsahin, Oddur Kjartansson, Alexander Gutkin, and Clara Rivera. Open-source Multi-speaker Corpora of the English Accents in the British Isles. In Proceedings of The 12th Language Resources and Evaluation Conference (LREC), pages 6532–6541, Marseille, France, May 2020. European Language Resources Association (ELRA). ISBN 979-10-95546-34-4. URL https://www.aclweb.org/anthology/2020.lrec-1.804.
  16. 16.Paul Ekman and Wallace V Friesen. Constants across cultures in the face and emotion. Journal of personality and social psychology, 17(2):124, 1971.
  17. 17.Jinlan Fu, See-Kiong Ng, Zhengbao Jiang, and Pengfei Liu. Gptscore: Evaluate as you desire. arXiv preprint arXiv:2302.04166, 2023.
  18. 18.Yuan Gong, Sameer Khurana, Leonid Karlinsky, and James Glass. Whisper-at: Noise-robust automatic speech recognizers are also strong general audio event taggers. arXiv preprint arXiv:2307.03183, 2023.
  19. 19.Yuan Gong, Alexander H Liu, Hongyin Luo, Leonid Karlinsky, and James Glass. Joint audio and speech understanding. In 2023 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), pages 1–8. IEEE, 2023.
  20. 20.Haorui He, Zengqiang Shang, Chaoren Wang, Xuyuan Li, Yicheng Gu, Hua Hua, Liwei Liu, Chen Yang, Jiaqi Li, Peiyang Shi, et al. Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation. arXiv preprint arXiv:2407.05361, 2024.
  21. 21.Dan Hendrycks and Kevin Gimpel. Gaussian error linear units (gelus). arXiv preprint arXiv:1606.08415, 2016.
  22. 22.Jonathan Ho and Tim Salimans. Classifier-free diffusion guidance. arXiv preprint arXiv:2207.12598, 2022.
  23. 23.Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. Lora: Low-rank adaptation of large language models. arXiv preprint arXiv:2106.09685, 2021.
  24. 24.Shujie Hu, Long Zhou, Shujie Liu, Sanyuan Chen, Hongkun Hao, Jing Pan, Xunying Liu, Jinyu Li, Sunit Sivasankaran, Linquan Liu, et al. Wavllm: Towards robust and adaptive speech large language model. arXiv preprint arXiv:2404.00656, 2024.
  25. 25.Jesin James, Li Tian, and Catherine Inez Watson. An Open Source Emotional Speech Corpus for Human Robot Interaction Applications. In Proc. Interspeech 2018, pages 2768–2772, 2018. doi: 10.21437/Interspeech.2018-1349.
  26. 26.Zeqian Ju, Yuancheng Wang, Kai Shen, Xu Tan, Detai Xin, Dongchao Yang, Yanqing Liu, Yichong Leng, Kaitao Song, Siliang Tang, et al. Naturalspeech 3: Zero-shot speech synthesis with factorized codec and diffusion models. arXiv preprint arXiv:2403.03100, 2024.
  27. 27.Jacob Kahn, Morgane Riviere, Weiyi Zheng, Evgeny Kharitonov, Qiantong Xu, Pierre-Emmanuel Mazaré, Julien Karadayi, Vitaliy Liptchinsky, Ronan Collobert, Christian Fuegen, et al. Libri-light: A benchmark for asr with limited or no supervision. In ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 7669–7673. IEEE, 2020.
  28. 28.Chris Dongjoo Kim, Byeongchang Kim, Hyunmin Lee, and Gunhee Kim. Audiocaps: Generating captions for audios in the wild. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 119–132, 2019.
  29. 29.Jaehyeon Kim, Keon Lee, Seungjun Chung, and Jaewoong Cho. Clam-tts: Improving neural codec language model for zero-shot text-to-speech. arXiv preprint arXiv:2404.02781, 2024.
  30. 30.Mateusz Łajszczak, Guillermo Cámbara, Yang Li, Fatih Beyhan, Arent van Korlaar, Fan Yang, Arnaud Joly, Álvaro Martín-Cortinas, Ammar Abbas, Adam Michalski, et al. Base tts: Lessons from building a billion-parameter text-to-speech model on 100k hours of data. arXiv preprint arXiv:2402.08093, 2024.
  31. 31.Chia-Hsuan Li, Szu-Lin Wu, Chi-Liang Liu, and Hung-yi Lee. Spoken squad: A study of mitigating the impact of speech recognition errors on listening comprehension. arXiv preprint arXiv:1804.00320, 2018.
  32. 32.Chin-Yew Lin. ROUGE: A package for automatic evaluation of summaries. In Text Summarization Branches Out, pages 74–81, Barcelona, Spain, July 2004. Association for Computational Linguistics. URL https://aclanthology.org/W04-1013.
  33. 33.Guan-Ting Lin, Cheng-Han Chiang, and Hung-yi Lee. Advancing large language models to capture varied speaking styles and respond properly in spoken conversations. arXiv preprint arXiv:2402.12786, 2024.
  34. 34.Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual instruction tuning. Advances in neural information processing systems, 36, 2024.
  35. 35.Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. G-eval: NLG evaluation using gpt-4 with better human alignment. In Houda Bouamor, Juan Pino, and Kalika Bali, editors, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 2511–2522, Singapore, December 2023. Association for Computational Linguistics. doi: 10.18653/v1/2023.emnlp-main.153. URL https://aclanthology.org/2023.emnlp-main.153.
  36. 36.Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101, 2017.
  37. 37.Reza Lotfian and Carlos Busso. Building naturalistic emotionally balanced speech corpus by retrieving emotional speech from existing podcast recordings. IEEE Transactions on Affective Computing, 10(4):471–483, 2019. doi: 10.1109/TAFFC.2017.2736999.
  38. 38.Ziyang Ma, Zhisheng Zheng, Jiaxin Ye, Jinchao Li, Zhifu Gao, Shiliang Zhang, and Xie Chen. emotion2vec: Self-supervised pre-training for speech emotion representation. arXiv preprint arXiv:2312.15185, 2023.
  39. 39.Gary McKeown, Michel F Valstar, Roderick Cowie, and Maja Pantic. The semaine corpus of emotionally coloured character interactions. In 2010 IEEE international conference on multimedia and expo, pages 1079–1084. IEEE, 2010.
  40. 40.Eliya Nachmani, Alon Levkovitch, Roy Hirsch, Julian Salazar, Chulayuth Asawaroengchai, Soroosh Mariooryad, Ehud Rivlin, RJ Skerry-Ryan, and Michelle Tadmor Ramanovich. Spoken question answering and speech continuation using spectrogram-powered llm. In The Twelfth International Conference on Learning Representations, 2023.
  41. 41.Clifford Nass and Scott Brave. Wired for Speech: How Voice Activates and Advances the Human-Computer Relationship. The MIT Press, 2005. ISBN 0262140926.
  42. 42.Jekaterina Novikova, Ondřej Dušek, Amanda Cercas Curry, and Verena Rieser. Why we need new evaluation metrics for nlg. arXiv preprint arXiv:1707.06875, 2017.
  43. 43.OpenAI. Chatgpt, 2022. https://openai.com/blog/chatgpt/.
  44. 44.OpenAI. Gpt-4o, 2024. https://openai.com/index/hello-gpt-4o/.
  45. 45.Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur. Librispeech: An asr corpus based on public domain audio books. In 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 5206–5210, 2015. doi: 10.1109/ICASSP.2015.7178964.
  46. 46.Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. Bleu: a method for automatic evaluation of machine translation. In Pierre Isabelle, Eugene Charniak, and Dekang Lin, editors, Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, pages 311–318, Philadelphia, Pennsylvania, USA, July 2002. Association for Computational Linguistics. doi: 10.3115/1073083.1073135. URL https://aclanthology.org/P02-1040.
  47. 47.Puyuan Peng, Po-Yao Huang, Daniel Li, Abdelrahman Mohamed, and David Harwath. Voicecraft: Zero-shot speech editing and text-to-speech in the wild. arXiv preprint arXiv:2403.16973, 2024.
  48. 48.Soujanya Poria, Devamanyu Hazarika, Navonil Majumder, Gautam Naik, Erik Cambria, and Rada Mihalcea. MELD: A multimodal multi-party dataset for emotion recognition in conversations. In Anna Korhonen, David Traum, and Lluís Màrquez, editors, Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 527–536, Florence, Italy, July 2019. Association for Computational Linguistics. doi: 10.18653/v1/P19-1050. URL https://aclanthology.org/P19-1050.
  49. 49.Sameer Pradhan, Ronald A. Cole, and Wayne H. Ward. My science tutor (MyST)–a large corpus of children’s conversational speech. In Nicoletta Calzolari, Min-Yen Kan, Veronique Hoste, Alessandro Lenci, Sakriani Sakti, and Nianwen Xue, editors, Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024), pages 12040–12045, Torino, Italia, May 2024. ELRA and ICCL. URL https://aclanthology.org/2024.lrec-main.1052.
  50. 50.Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever. Robust speech recognition via large-scale weak supervision. In International Conference on Machine Learning, pages 28492–28518. PMLR, 2023.
  51. 51.Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn. Direct preference optimization: Your language model is secretly a reward model. Advances in Neural Information Processing Systems, 36, 2024.
  52. 52.Kashfia Sailunaz and Reda Alhajj. Emotion and sentiment analysis from twitter text. Journal of computational science, 36:101003, 2019.
  53. 53.Guillaume Sanchez, Honglu Fan, Alexander Spangher, Elad Levi, Pawan Sasanka Ammanamanchi, and Stella Biderman. Stay on topic with classifier-free guidance. arXiv preprint arXiv:2306.17806, 2023.
  54. 54.Dan Su and Pascale Fung. Improving spoken question answering using contextualized word representation. In ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 8004–8008. IEEE, 2020.
  55. 55.Changli Tang, Wenyi Yu, Guangzhi Sun, Xianzhao Chen, Tian Tan, Wei Li, Lu Lu, Zejun MA, and Chao Zhang. SALMONN: Towards generic hearing abilities for large language models. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=14rn7HpKVk.
  56. 56.Gemini Team, Rohan Anil, Sebastian Borgeaud, Yonghui Wu, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, et al. Gemini: a family of highly capable multimodal models. arXiv preprint arXiv:2312.11805, 2023.
  57. 57.Gemma Team, Morgane Riviere, Shreya Pathak, Pier Giuseppe Sessa, Cassidy Hardin, Surya Bhupatiraju, Léonard Hussenot, Thomas Mesnard, Bobak Shahriari, Alexandre Ramé, et al. Gemma 2: Improving open language models at a practical size. arXiv preprint arXiv:2408.00118, 2024.
  58. 58.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023.
  59. 59.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023.
  60. 60.Bo-Hsiang Tseng, Sheng-Syun Shen, Hung-Yi Lee, and Lin-Shan Lee. Towards machine comprehension of spoken content: Initial toefl listening comprehension test by machine. arXiv preprint arXiv:1608.06378, 2016.
  61. 61.Chengyi Wang, Sanyuan Chen, Yu Wu, Ziqiang Zhang, Long Zhou, Shujie Liu, Zhuo Chen, Yanqing Liu, Huaming Wang, Jinyu Li, et al. Neural codec language models are zero-shot text to speech synthesizers. arXiv preprint arXiv:2301.02111, 2023.
  62. 62.Kaisiyuan Wang, Qianyi Wu, Linsen Song, Zhuoqian Yang, Wayne Wu, Chen Qian, Ran He, Yu Qiao, and Chen Change Loy. Mead: A large-scale audio-visual dataset for emotional talking-face generation. In ECCV, 2020.
  63. 63.Yijing Wu, SaiKrishna Rallabandi, Ravisutha Srinivasamurthy, Parag Pravin Dakle, Alolika Gon, and Preethi Raghavan. Heysquad: A spoken question answering dataset. arXiv preprint arXiv:2304.13689, 2023.
  64. 64.Hongfei Xue, Yuhao Liang, Bingshen Mu, Shiliang Zhang, Qian Chen, and Lei Xie. E-chat: Emotion-sensitive spoken dialogue system with large language models. arXiv preprint arXiv:2401.00475, 2023.
  65. 65.Junichi Yamagishi, Christophe Veaux, and Kirsten MacDonald. CSTR VCTK Corpus: English multi-speaker corpus for CSTR voice cloning toolkit, 2019.
  66. 66.An Yang, Baosong Yang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Zhou, Chengpeng Li, Chengyuan Li, Dayiheng Liu, Fei Huang, et al. Qwen2 technical report. arXiv preprint arXiv:2407.10671, 2024.
  67. 67.Chenyu You, Nuo Chen, Fenglin Liu, Shen Ge, Xian Wu, and Yuexian Zou. End-to-end spoken conversational question answering: Task, dataset and model. arXiv preprint arXiv:2204.14272, 2022.
  68. 68.Alex Young, Bei Chen, Chao Li, Chengen Huang, Ge Zhang, Guanwei Zhang, Heng Li, Jiangcheng Zhu, Jianqun Chen, Jing Chang, et al. Yi: Open foundation models by 01. ai. arXiv preprint arXiv:2403.04652, 2024.
  69. 69.Binbin Zhang, Hang Lv, Pengcheng Guo, Qijie Shao, Chao Yang, Lei Xie, Xin Xu, Hui Bu, Xiaoyu Chen, Chenchen Zeng, et al. Wenetspeech: A 10000+ hours multi-domain mandarin corpus for speech recognition. In ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 6182–6186. IEEE, 2022.
  70. 70.Dong Zhang, Shimin Li, Xin Zhang, Jun Zhan, Pengyu Wang, Yaqian Zhou, and Xipeng Qiu. SpeechGPT: Empowering large language models with intrinsic cross-modal conversational abilities. In Houda Bouamor, Juan Pino, and Kalika Bali, editors, Findings of the Association for Computational Linguistics: EMNLP 2023, pages 15757–15773, Singapore, December 2023. Association for Computational Linguistics. doi: 10.18653/v1/2023.findings-emnlp.1055. URL https://aclanthology.org/2023.findings-emnlp.1055.
  71. 71.Pan Zhang, Xiaoyi Dong Bin Wang, Yuhang Cao, Chao Xu, Linke Ouyang, Zhiyuan Zhao, Shuangrui Ding, Songyang Zhang, Haodong Duan, Hang Yan, et al. Internlm-xcomposer: A vision-language large model for advanced text-image comprehension and composition. arXiv preprint arXiv:2309.15112, 2023.
  72. 72.Tianyi Zhang*, Varsha Kishore*, Felix Wu*, Kilian Q. Weinberger, and Yoav Artzi. Bertscore: Evaluating text generation with bert. In International Conference on Learning Representations, 2020. URL https://openreview.net/forum?id=SkeHuCVFDr.
  73. 73.Xueyao Zhang, Liumeng Xue, Yicheng Gu, Yuancheng Wang, Jiaqi Li, Haorui He, Chaoren Wang, Ting Song, Xi Chen, Zihao Fang, Haopeng Chen, Junan Zhang, Tze Ying Tang, Lexiao Zou, Mingxuan Wang, Jun Han, Kai Chen, Haizhou Li, and Zhizheng Wu. Amphion: An open-source audio, music and speech generation toolkit. In IEEE Spoken Language Technology Workshop, SLT 2024, 2024.
  74. 74.Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al. Judging llm-as-a-judge with mt-bench and chatbot arena. Advances in Neural Information Processing Systems, 36, 2024.
  75. 75.Kun Zhou, Berrak Sisman, Rui Liu, and Haizhou Li. Emotional voice conversion: Theory, databases and esd. Speech Communication, 137:1–18, 2022. ISSN 0167-6393. doi: https://doi.org/10.1016/j.specom.2021.11.006. URL https://www.sciencedirect.com/science/article/pii/S0167639321001308.
  76. 76.Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny. Minigpt-4: Enhancing vision-language understanding with advanced large language models. arXiv preprint arXiv:2304.10592, 2023.

Citation

MLA
Ao, J., et al. “SD-Eval: A Benchmark Dataset for Spoken Dialogue Understanding Beyond Words”. arXiv, 2024, http://arxiv.org/abs/2406.13340v2.
APA
Ao, J., Wang, Y., Tian, X., Chen, D., Zhang, J., Lu, L., Wang, Y., Li, H., & Wu, Z. (2024). SD-Eval: A Benchmark Dataset for Spoken Dialogue Understanding Beyond Words. arXiv. http://arxiv.org/abs/2406.13340v2
Chicago
Ao, J., Y. Wang, X. Tian, et al. 2024. “SD-Eval: A Benchmark Dataset for Spoken Dialogue Understanding Beyond Words”. arXiv. http://arxiv.org/abs/2406.13340v2.
Harvard
Ao, J. et al. (2024) “SD-Eval: A Benchmark Dataset for Spoken Dialogue Understanding Beyond Words”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2406.13340v2.
Vancouver
1. Ao J, Wang Y, Tian X, Chen D, Zhang J, Lu L, Wang Y, Li H, Wu Z (2024) SD-Eval: A Benchmark Dataset for Spoken Dialogue Understanding Beyond Words. arXiv

BibTeX

@article{ao2024eval,
  title = {SD-Eval: A Benchmark Dataset for Spoken Dialogue Understanding Beyond Words},
  author = {Ao, Junyi and Wang, Yuancheng and Tian, Xiaohai and Chen, Dekun and Zhang, Jun and Lu, Lu and Wang, Yuxuan and Li, Haizhou and Wu, Zhizheng},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2406.13340v2},
  eprint = {2406.13340}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors