Built independently by an author, for readers. Read the story and support ChapterPal

keyword

speech emotion captioning

Speech emotion captioning is an artificial intelligence and speech processing task that automatically generates descriptive natural language text to explain the emotional states and nuances expressed in human speech. Unlike traditional speech emotion recognition, which assigns vocal audio to predetermined discrete categories or numerical values, speech emotion captioning produces open-ended, contextualized descriptions of vocal tone, mood, intensity, and subtle expressive shifts. The task typically couples audio encoders with generative language models to translate complex acoustic cues and speech characteristics into coherent, human-readable sentences that provide a richer and more flexible representation of human emotion.

1 item

SECap: Speech Emotion Captioning with Large Language Model

SECap: Speech Emotion Captioning with Large Language Model

Yaoxun Xu, Hangting Chen, Jianwei Yu, Qiaochu Huang, Zhiyong Wu, Shi-Xiong Zhang, Guangzhi Li, Yi Luo, Rongzhi Gu

OrganizationsTencentThe Chinese University of Hong KongTsinghua University

Why you should read this

Proposes SECap, a framework that integrates HuBERT audio representations with LLaMA through a Q-Former to generate rich natural language descriptions of complex speech emotions rather than relying on discrete classification labels.

Speech emotions are crucial in human communication and are extensively used in fields like speech synthesis and natural language understanding. Most prior studies, such as speech emotion recognition, have categorized speech emotions into a fixed set of classes. Yet, emotions expressed in human speech are often complex, and categorizing them into predefined groups can be insufficient to adequately represent speech emotions. On the contrary, describing speech emotions directly by means of natural language may be a more effective approach. Regrettably, there are not many studies available that have focused on this direction. Therefore, this paper proposes a speech emotion captioning framework named SECap, aiming at effectively describing speech emotions using natural language. Owing to the impressive capabilities of large language models in language comprehension and text generation, SECap employs LLaMA as the text decoder to allow the production of coherent speech emotion captions. In addition, SECap leverages HuBERT as the audio encoder to extract general speech features and Q-Former as the Bridge-Net to provide LLaMA with emotion-related speech features. To accomplish this, Q-Former utilizes mutual information learning to disentangle emotion-related speech features and speech contents, while implementing contrastive learning to extract more emotion-related speech features. The results of objective and subjective evaluations demonstrate that: 1) the SECap framework outperforms the HTSAT-BART baseline in all objective evaluations; 2) SECap can generate high-quality speech emotion captions that attain performance on par with human annotators in subjective mean opinion score tests.

Added

2026-09-26