SECap: Speech Emotion Captioning with Large Language Model
Yaoxun XuHangting ChenJianwei YuQiaochu HuangZhiyong WuShi-Xiong ZhangGuangzhi LiYi LuoRongzhi Gu
Proposes SECap, a framework that integrates HuBERT audio representations with LLaMA through a Q-Former to generate rich natural language descriptions of complex speech emotions rather than relying on discrete classification labels.
Interpreting vocal emotion is critical for effective human-machine communication, yet traditional automated systems rely almost entirely on rigid classification categories such as anger, fear, or happiness. Real-world human emotions are multifaceted and nuanced, frequently exhibiting mixed states and varying intensities that single-word labels fail to capture. The article addresses this limitation by introducing the task of Speech Emotion Captioning, which uses natural language sentences instead of fixed discrete labels to describe acoustic emotional states.
The main objective of the article is to propose and evaluate SECap, an end-to-end framework designed to generate descriptive, fluent, and human-like natural language captions of speech emotions directly from audio. It demonstrates how integrating specialized feature extraction mechanisms with large language models can produce descriptive assessments that match human quality.
To achieve this, the authors designed a three-part architecture: a speech encoder (HuBERT) to extract audio representations, a bridging network (Q-Former) to isolate emotion-specific acoustic features from spoken content and compress the data, and a large language model (LLaMA) to generate coherent text descriptions. The framework employs a two-stage training strategy combining mutual information learning to remove spoken content influence with contrastive learning to emphasize emotional cues. The system was trained and evaluated on EMOSpeech, a 41.6-hour dataset comprising 30,526 annotated Chinese speech recordings, and benchmarked against standard audio captioning models using objective similarity metrics and subjective human evaluation panels.
The findings confirm the superiority of natural language descriptions over traditional labeling. SECap outperformed the HTSAT-BART baseline across all objective evaluation metrics, including word matching and sentence similarity benchmarks. In subjective assessments, SECap achieved a Mean Opinion Score of 3.77 out of 5.0, surpassing standard single-word emotion classification models (3.39) and performing on par with ground-truth human annotations (3.85). Ablation studies showed that combining content-disentanglement and contrastive learning boosted similarity performance by about 6.9%, while actively tuning the bridging network alongside the language model prevented a 19.5% drop in performance.
These results show that large language models can effectively describe emotional nuances in speech, providing a richer and more practical tool for downstream applications such as conversational agents, sentiment analysis, and speech synthesis. Natural language descriptions reduce the ambiguity inherent in forced-choice emotion labeling and align better with human emotional perception. For operational implementation, organizations developing advanced voice interfaces can adopt descriptive speech emotion frameworks to capture complex speaker states more effectively than discrete classifiers.
Decision-makers should note certain limitations. The evaluations were conducted on a single Mandarin dataset containing only seven speakers, and the full pipeline requires significant computational resources across its multi-billion-parameter architecture. While confidence in the model's performance on the evaluated dataset is high, further validation on broader, multi-speaker, and multilingual datasets is recommended before enterprise-scale deployment.
- Paper: Pengi: An Audio Language Model for Audio Tasks, Soham Deshmukh et al. (2023). This paper establishes the paradigm of framing unified audio understanding and captioning as language model text-generation tasks, which SECap directly builds upon for speech emotion captioning.
- Paper: MISA: Modality-Invariant and -Specific Representations for Multimodal Sentiment Analysis, Devamanyu Hazarika et al. (2020). This work introduces techniques for disentangling shared invariant representations from specific feature subspaces in multimodal sentiment analysis, providing conceptual foundations for SECap's disentanglement of emotion cues from spoken content.
- Paper: Multimodal Transformer for Unaligned Multimodal Language Sequences, Yao-Hung Hubert Tsai et al. (2019). This paper provides foundational crossmodal attention architectures for processing unaligned multimodal language and acoustic sequences, which are essential for connecting speech encoders to language decoders.
- Paper: MELD: A Multimodal Multi-Party Dataset for Emotion Recognition in Conversations, Soujanya Poria et al. (2018). This study introduces the multi-modal emotional dialogue benchmark (MELD) and highlights the limitations of discrete emotion classification that SECap seeks to overcome through natural language captioning.
- Paper: Multimodal Language Analysis in the Wild: CMU-MOSEI Dataset and Interpretable Dynamic Fusion Graph, Amir Zadeh et al. (2018). This work establishes large-scale multimodal sentiment and emotion intensity benchmarks (CMU-MOSEI), framing the complex interplay between acoustics and emotional expressions addressed in SECap.
- Paper: SD-Eval: A Benchmark Dataset for Spoken Dialogue Understanding Beyond Words, Junyi Ao et al. (2024). SD-Eval evaluates spoken dialogue systems on understanding non-verbal acoustic cues such as vocal emotion, providing an ideal benchmark for downstream applications using descriptive speech emotion representations like SECap.
- Paper: BRIGHTER: BRIdging the Gap in Human-Annotated Textual Emotion Recognition Datasets for 28 Languages, Shamsuddeen Hassan Muhammad et al. (2025). BRIGHTER extends multilingual and multi-label emotion recognition evaluation across 28 diverse languages, offering a valuable reference for addressing SECap's limitation to single-language Mandarin corpora.
- Paper: Can Large Language Models be Good Emotional Supporter? Mitigating Preference Bias on Emotional Support Conversation, Dongjin Kang et al. (2024). This paper examines how large language models deliver emotional support in conversations, representing a direct downstream application for fine-grained speech emotion descriptions.
- Paper: Qwen3.5-Omni Technical Report, Qwen Team (2026). This technical report demonstrates unified omnimodal foundation models combining speech perception, reasoning, and synthesis, scaling up the end-to-end audio-language integration demonstrated by SECap.
