Built independently by an author, for readers. Read the story and support ChapterPal

topic

speech emotion recognition

Speech emotion recognition is a subfield of artificial intelligence and audio signal processing that focuses on automatically identifying and classifying human emotional states from spoken audio. By analyzing acoustic and prosodic properties such as pitch, energy, vocal timbre, rhythm, and spectral characteristics, computational models extract distinctive features from speech signals to distinguish emotional states like anger, joy, sadness, fear, or neutrality. These features are evaluated using machine learning algorithms and neural network architectures to infer the speaker affective condition regardless of semantic content. The technology is widely utilized in affective computing, human-computer interaction, intelligent call centers, healthcare diagnostics, and interactive voice response systems to enable more natural and context-aware communication between humans and machines.

1 item

SECap: Speech Emotion Captioning with Large Language Model

SECap: Speech Emotion Captioning with Large Language Model

Yaoxun Xu, Hangting Chen, Jianwei Yu, Qiaochu Huang, Zhiyong Wu, Shi-Xiong Zhang, Guangzhi Li, Yi Luo, Rongzhi Gu

OrganizationsTencentThe Chinese University of Hong KongTsinghua University

Why you should read this

Proposes SECap, a framework that integrates HuBERT audio representations with LLaMA through a Q-Former to generate rich natural language descriptions of complex speech emotions rather than relying on discrete classification labels.

Speech emotions are crucial in human communication and are extensively used in fields like speech synthesis and natural language understanding. Most prior studies, such as speech emotion recognition, have categorized speech emotions into a fixed set of classes. Yet, emotions expressed in human speech are often complex, and categorizing them into predefined groups can be insufficient to adequately represent speech emotions. On the contrary, describing speech emotions directly by means of natural language may be a more effective approach. Regrettably, there are not many studies available that have focused on this direction. Therefore, this paper proposes a speech emotion captioning framework named SECap, aiming at effectively describing speech emotions using natural language. Owing to the impressive capabilities of large language models in language comprehension and text generation, SECap employs LLaMA as the text decoder to allow the production of coherent speech emotion captions. In addition, SECap leverages HuBERT as the audio encoder to extract general speech features and Q-Former as the Bridge-Net to provide LLaMA with emotion-related speech features. To accomplish this, Q-Former utilizes mutual information learning to disentangle emotion-related speech features and speech contents, while implementing contrastive learning to extract more emotion-related speech features. The results of objective and subjective evaluations demonstrate that: 1) the SECap framework outperforms the HTSAT-BART baseline in all objective evaluations; 2) SECap can generate high-quality speech emotion captions that attain performance on par with human annotators in subjective mean opinion score tests.

Added

2026-09-26