SD-Eval: A Benchmark Dataset for Spoken Dialogue Understanding Beyond Words
Junyi AoYuancheng WangXiaohai TianDekun ChenJun ZhangLu LuYuxuan WangHaizhou LiZhizheng Wu
Introduces SD-Eval, an open-source benchmark and training resource designed to evaluate how effectively speech language models integrate speaker emotion, accent, age, and background sounds into dialogue generation.
Spoken communication conveys rich contextual cues beyond literal words, including vocal emotions, accents, speaker age, and environmental background noises. Although modern speech-enabled artificial intelligence systems can transcribe words and perceive audio features, they frequently fail to generate context-aware, empathetic responses in real conversations. This limitation stems from a lack of standard task definitions, open-source datasets, and specialized evaluation frameworks designed to assess conversational responses conditioned on non-verbal audio cues.
The article introduces SD-Eval, a benchmark dataset designed to evaluate spoken dialogue understanding and response generation across multiple dimensions beyond text alone. The objective is to establish standardized resources and metrics that promote the development of intelligent dialogue systems capable of adapting responses to vocal and environmental nuances.
The researchers constructed the SD-Eval benchmark using 7,303 speech utterances totaling 8.76 hours across four evaluation sub-tasks: emotion, accent, age, and environmental sounds. The data was curated from eight public speech datasets, incorporating both real recordings and targeted synthetic speech. To support baseline model development, the researchers assembled a corresponding training dataset containing 724,400 utterances totaling 1,052.72 hours across eleven public datasets. They evaluated baseline systems—including text-cascaded models, an end-to-end speech large language model, and several open-source audio language models—using traditional text-matching metrics, advanced large language model judges, and human evaluations across 200 sampled conversations.
The findings demonstrate that the end-to-end speech model, which ingests raw audio representations directly, consistently outperformed traditional cascaded systems that rely solely on automated text transcription. For example, on the emotion subset, the end-to-end model achieved a response quality score of 5.30 compared to 4.47 for the cascaded baseline when rated by an advanced model judge, reflecting its ability to pick up non-verbal cues directly from speech. Providing explicit, high-quality ground-truth contextual labels produced the highest overall performance, demonstrating that input quality significantly influences conversational appropriateness. Furthermore, off-the-shelf open-source models performed poorly on these nuanced conversational tasks unless specifically operated in dedicated voice-chat modes. Methodologically, automated evaluations conducted using advanced large language models aligned substantially better with human judgments than traditional metrics like BLEU or ROUGE, with the top model judge achieving a strong overall correlation of 0.613 compared to 0.208 for BLEU-4.
These results indicate that conversational systems relying solely on speech-to-text pipelines risk misinterpreting user intent by ignoring paralinguistic and ambient context. Transitioning toward end-to-end audio processing or incorporating explicit acoustic feature recognition is essential for deploying conversational artificial intelligence in customer service, healthcare, and voice assistance where empathy and situational awareness are critical. Additionally, organizations can adopt large language model judges as reliable, scalable evaluation proxies, substantially reducing the high financial and operational costs associated with manual human quality reviews.
Stakeholders developing voice-based interactive systems should transition from pure text-cascaded architectures to end-to-end multi-modal models trained explicitly on diverse acoustic contexts. Development teams should expand training data to include broader demographic characteristics, explore feature-disentanglement methods, and adopt model-based evaluation frameworks to monitor response quality. Readers should note that SD-Eval is currently constrained to single-turn, speech-to-text dialogues and does not evaluate speaker gender or multi-turn speech-to-speech interactions, representing key areas for future research.
- Paper: MELD: A Multimodal Multi-Party Dataset for Emotion Recognition in Conversations, Soujanya Poria et al. (2018). Provides a core multimodal conversational dataset and baseline methodologies for emotion recognition in dialogue that motivate paralinguistic modeling in spoken conversation.
- Paper: Multimodal Language Analysis in the Wild: CMU-MOSEI Dataset and Interpretable Dynamic Fusion Graph, Amir Zadeh et al. (2018). Establishes standard multimodal sentiment and emotion representations across acoustic and linguistic modalities, foundational for paralinguistic spoken dialogue understanding.
- Paper: DailyDialog: A Manually Labelled Multi-turn Dialogue Dataset, Yanran Li et al. (2017). Introduces high-quality multi-turn dialogue emotion and intent annotation paradigms that prefigure dialogue understanding beyond literal content.
- Paper: How NOT To Evaluate Your Dialogue System: An Empirical Study of Unsupervised Evaluation Metrics for Dialogue Response Generation, Chia-Wei Liu et al. (2016). Highlights the critical flaws of word-overlap metrics like BLEU and ROUGE in evaluating open-ended dialogue response generation, justifying the need for LLM-based evaluation metrics.
- Paper: Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena, Lianmin Zheng et al. (2023). Demonstrates the efficacy and methodology of using LLM-as-a-judge for evaluating open-ended dialogue responses, which SD-Eval adopts for multi-dimensional evaluation.
- Paper: WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing, Sanyuan Chen et al. (2021). Presents foundational self-supervised acoustic representations robust to noise and paralinguistic variation used in modern speech-LLM dialogue architectures.
- Paper: G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment, Yang Liu et al. (2023). Develops form-filling and chain-of-thought LLM-based NLG evaluation frameworks that provide the groundwork for LLM scoring of spoken dialogue generation.
- Paper: SECap: Speech Emotion Captioning with Large Language Model, Yaoxun Xu et al. (2024). Extends spoken paralinguistic understanding from dialogue response generation to open-ended natural language captioning of complex speech emotions with LLMs.
- Paper: MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark, S. Sakshi et al. (2025). Expands comprehensive audio and speech evaluation into a massive multi-task reasoning and expert-level audio understanding benchmark.
- Paper: From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judge, Dawei Li et al. (2025). Synthesizes and formalizes the LLM-as-a-judge evaluation techniques utilized and validated in multi-turn spoken dialogue benchmarks like SD-Eval.
- Paper: Open-Source Conversational AI with SpeechBrain 1.0, Mirco Ravanelli et al. (2024). Implements an open-source conversational AI framework integrating speech and text LLMs that operationalizes end-to-end spoken dialogue systems.
