Built independently by an author, for readers. Read the story and support ChapterPal

keyword

speech-caption contrastive learning

Speech-caption contrastive learning is a multimodal machine learning technique that aligns spoken audio signals with descriptive natural language texts in a shared embedding space. In this approach, neural networks process speech recordings and textual descriptions concurrently, utilizing an objective function that pulls representations of corresponding speech-text pairs closer together while pushing unrelated pairs further apart. By learning to associate complex acoustic, semantic, and paralinguistic cues with descriptive vocabulary, this method bridges audio representations with natural language decoders, supporting downstream tasks such as automated speech captioning, speech emotion analysis, and cross-modal audio-text retrieval.

1 item

SECap: Speech Emotion Captioning with Large Language Model

SECap: Speech Emotion Captioning with Large Language Model

Yaoxun Xu, Hangting Chen, Jianwei Yu, Qiaochu Huang, Zhiyong Wu, Shi-Xiong Zhang, Guangzhi Li, Yi Luo, Rongzhi Gu

OrganizationsTencentThe Chinese University of Hong KongTsinghua University

Why you should read this

Proposes SECap, a framework that integrates HuBERT audio representations with LLaMA through a Q-Former to generate rich natural language descriptions of complex speech emotions rather than relying on discrete classification labels.

Speech emotions are crucial in human communication and are extensively used in fields like speech synthesis and natural language understanding. Most prior studies, such as speech emotion recognition, have categorized speech emotions into a fixed set of classes. Yet, emotions expressed in human speech are often complex, and categorizing them into predefined groups can be insufficient to adequately represent speech emotions. On the contrary, describing speech emotions directly by means of natural language may be a more effective approach. Regrettably, there are not many studies available that have focused on this direction. Therefore, this paper proposes a speech emotion captioning framework named SECap, aiming at effectively describing speech emotions using natural language. Owing to the impressive capabilities of large language models in language comprehension and text generation, SECap employs LLaMA as the text decoder to allow the production of coherent speech emotion captions. In addition, SECap leverages HuBERT as the audio encoder to extract general speech features and Q-Former as the Bridge-Net to provide LLaMA with emotion-related speech features. To accomplish this, Q-Former utilizes mutual information learning to disentangle emotion-related speech features and speech contents, while implementing contrastive learning to extract more emotion-related speech features. The results of objective and subjective evaluations demonstrate that: 1) the SECap framework outperforms the HTSAT-BART baseline in all objective evaluations; 2) SECap can generate high-quality speech emotion captions that attain performance on par with human annotators in subjective mean opinion score tests.

Added

2026-09-26