The SD-Eval benchmark dataset is a speech evaluation benchmark designed to assess how well artificial intelligence models understand spoken dialogue and generate conversational responses based on audio features beyond written text. It specifically measures a model's ability to interpret paralinguistic and environmental cues across four primary dimensions: speaker emotion, accent, speaker age, and background sounds. By evaluating conversational systems on these non-lexical audio factors, the dataset provides standardized test material to analyze and improve the contextual awareness and empathy of spoken dialogue models.