SECap: Speech Emotion Captioning with Large Language Model

Yaoxun XuHangting ChenJianwei YuQiaochu HuangZhiyong WuShi-Xiong ZhangGuangzhi LiYi LuoRongzhi Gu

article2024AAAI84 citations

Proposes SECap, a framework that integrates HuBERT audio representations with LLaMA through a Q-Former to generate rich natural language descriptions of complex speech emotions rather than relying on discrete classification labels.

Listen

Interpreting vocal emotion is critical for effective human-machine communication, yet traditional automated systems rely almost entirely on rigid classification categories such as anger, fear, or happiness. Real-world human emotions are multifaceted and nuanced, frequently exhibiting mixed states and varying intensities that single-word labels fail to capture. The article addresses this limitation by introducing the task of Speech Emotion Captioning, which uses natural language sentences instead of fixed discrete labels to describe acoustic emotional states.

The main objective of the article is to propose and evaluate SECap, an end-to-end framework designed to generate descriptive, fluent, and human-like natural language captions of speech emotions directly from audio. It demonstrates how integrating specialized feature extraction mechanisms with large language models can produce descriptive assessments that match human quality.

To achieve this, the authors designed a three-part architecture: a speech encoder (HuBERT) to extract audio representations, a bridging network (Q-Former) to isolate emotion-specific acoustic features from spoken content and compress the data, and a large language model (LLaMA) to generate coherent text descriptions. The framework employs a two-stage training strategy combining mutual information learning to remove spoken content influence with contrastive learning to emphasize emotional cues. The system was trained and evaluated on EMOSpeech, a 41.6-hour dataset comprising 30,526 annotated Chinese speech recordings, and benchmarked against standard audio captioning models using objective similarity metrics and subjective human evaluation panels.

The findings confirm the superiority of natural language descriptions over traditional labeling. SECap outperformed the HTSAT-BART baseline across all objective evaluation metrics, including word matching and sentence similarity benchmarks. In subjective assessments, SECap achieved a Mean Opinion Score of 3.77 out of 5.0, surpassing standard single-word emotion classification models (3.39) and performing on par with ground-truth human annotations (3.85). Ablation studies showed that combining content-disentanglement and contrastive learning boosted similarity performance by about 6.9%, while actively tuning the bridging network alongside the language model prevented a 19.5% drop in performance.

These results show that large language models can effectively describe emotional nuances in speech, providing a richer and more practical tool for downstream applications such as conversational agents, sentiment analysis, and speech synthesis. Natural language descriptions reduce the ambiguity inherent in forced-choice emotion labeling and align better with human emotional perception. For operational implementation, organizations developing advanced voice interfaces can adopt descriptive speech emotion frameworks to capture complex speaker states more effectively than discrete classifiers.

Decision-makers should note certain limitations. The evaluations were conducted on a single Mandarin dataset containing only seven speakers, and the full pipeline requires significant computational resources across its multi-billion-parameter architecture. While confidence in the model's performance on the evaluated dataset is high, further validation on broader, multi-speaker, and multilingual datasets is recommended before enterprise-scale deployment.

Cover for SECap: Speech Emotion Captioning with Large Language Model

Abstract

Speech emotions are crucial in human communication and are extensively used in fields like speech synthesis and natural language understanding. Most prior studies, such as speech emotion recognition, have categorized speech emotions into a fixed set of classes. Yet, emotions expressed in human speech are often complex, and categorizing them into predefined groups can be insufficient to adequately represent speech emotions. On the contrary, describing speech emotions directly by means of natural language may be a more effective approach. Regrettably, there are not many studies available that have focused on this direction. Therefore, this paper proposes a speech emotion captioning framework named SECap, aiming at effectively describing speech emotions using natural language. Owing to the impressive capabilities of large language models in language comprehension and text generation, SECap employs LLaMA as the text decoder to allow the production of coherent speech emotion captions. In addition, SECap leverages HuBERT as the audio encoder to extract general speech features and Q-Former as the Bridge-Net to provide LLaMA with emotion-related speech features. To accomplish this, Q-Former utilizes mutual information learning to disentangle emotion-related speech features and speech contents, while implementing contrastive learning to extract more emotion-related speech features. The results of objective and subjective evaluations demonstrate that: 1) the SECap framework outperforms the HTSAT-BART baseline in all objective evaluations; 2) SECap can generate high-quality speech emotion captions that attain performance on par with human annotators in subjective mean opinion score tests.

Table of Contents

  • Introduction
  • Related Work
  • Speech Emotion Recognition
  • Large Language Model
  • Automated Audio Captioning
  • Method
  • Model Architecture
  • Obtain Emotion-Related Representations
  • Training Process
  • Dataset
  • Evaluation Metric
  • Objective Evaluation
  • Subjective Evaluation
  • Results and Analysis
  • Experiment Setup
  • Performance Analysis
  • Ablation Study on Different Model Components
  • Comparison of Training Methods
  • Conclusion
  • Acknowledgements
  • References

Knowls

  1. Knowl 1 — Speech Emotion Captioning (SEC) Task

    definition

    Speech Emotion Captioning (SEC) is the task of generating natural language descriptive sentences to characterize the emotional state, emotional intensity, acoustic fluctuations, and vocal nuances (such as volume and speech rate) of a speaker directly from an input speech utterance.

    In contrast to traditional Speech Emotion Recognition (SER), which categorizes speech into a single discrete class label (e.g., "angry", "happy"), SEC produces descriptive text capable of capturing multifaceted, co-occurring affective states and fine-grained emotional dynamics within an utterance.

  2. Knowl 2 — SECap Framework Architecture

    model/method

    The SECap framework is an encoder-decoder architecture designed for Speech Emotion Captioning that connects an audio encoder, a bridging network, and a large language model (LLM) text decoder:

    1. Audio Encoder: A pre-trained HuBERT (Hidden-Unit BERT) model extracts frame-level speech embeddings S∈Rns×Ts×dsS \in \mathbb{R}^{n_s \times T_s \times d_s} from raw speech input, where nsn_s is batch size, TsT_s is the number of time steps, and dsd_s is the feature dimension.
    2. Bridge-Net (Q-Former): A Querying Transformer (Q-Former) with learnable query vectors compresses the variable-length speech embeddings SS into a fixed-length representation Qe∈Rns×nq×dqQ_e \in \mathbb{R}^{n_s \times n_q \times d_q} (nqn_q queries of dimension dqd_q). A linear projection layer then maps QeQ_e into the text decoder's embedding space as prefix embedding L-Embedding\text{L-Embedding}.
    3. Text Decoder (LLaMA): A large language model (LLaMA, fine-tuned on Chinese text) acts as the autoregressive text decoder. The input sequence to LLaMA is constructed as: [BOS]+L-Embedding+Prompt_Emb+Caption\text{[BOS]} + \text{L-Embedding} + \text{Prompt\_Emb} + \text{Caption} where [BOS]\text{[BOS]} is the beginning-of-sequence token, and Prompt_Emb\text{Prompt\_Emb} is an instruction token sequence (e.g., instructing the model to describe the speaker's emotion in a single sentence) that constrains LLaMA's text generation.
  3. Knowl 3 — Q-Former Feature Compression and Attention Mechanism in SECap

    model/method

    In the SECap architecture, the Q-Former compresses variable-length frame-level HuBERT representations into fixed-length emotion representations using stacked self-attention and cross-attention modules.

    Let q∈Rnq×dqq \in \mathbb{R}^{n_q \times d_q} denote nqn_q learnable Q-queries of dimension dqd_q, and let S∈Rns×Ts×dsS \in \mathbb{R}^{n_s \times T_s \times d_s} denote the HuBERT speech embedding with batch size nsn_s, time steps TsT_s, and feature dimension dsd_s.

    First, self-attention is computed among the Q-queries: Aself=softmax(qWqself(qWkself)Tdk)qWvselfA_{\text{self}} = \text{softmax}\left( \frac{q W_{q\text{self}} (q W_{k\text{self}})^T}{\sqrt{d_k}} \right) q W_{v\text{self}} where Wqself∈Rdq×dkW_{q\text{self}} \in \mathbb{R}^{d_q \times d_k}, Wkself∈Rdq×dkW_{k\text{self}} \in \mathbb{R}^{d_q \times d_k}, and Wvself∈Rdq×dvW_{v\text{self}} \in \mathbb{R}^{d_q \times d_v} are learnable projection matrices, and dk,dvd_k, d_v are key and value dimensions.

    Next, the self-attention output Aself∈Rnq×dvA_{\text{self}} \in \mathbb{R}^{n_q \times d_v} serves as queries, while the speech embedding SS serves as keys and values in cross-attention: Across=softmax(AselfWq(SWk)Tdk)SWvA_{\text{cross}} = \text{softmax}\left( \frac{A_{\text{self}} W_q (S W_k)^T}{\sqrt{d_k}} \right) S W_v where Wq∈Rdv×dkW_q \in \mathbb{R}^{d_v \times d_k}, Wk∈Rds×dkW_k \in \mathbb{R}^{d_s \times d_k}, and Wv∈Rds×dvW_v \in \mathbb{R}^{d_s \times d_v}.

    The resulting Q-Embedding Qe∈Rns×nq×dqQ_e \in \mathbb{R}^{n_s \times n_q \times d_q} maintains a fixed length nqn_q independent of speech duration TsT_s.

  4. Knowl 4 — Speech-Transcription Mutual Information Learning (STMIL)

    model/method

    Speech-Transcription Mutual Information Learning (STMIL) is an optimization objective used in SECap's Bridge-Net to disentangle emotional acoustic features from spoken linguistic content. Spoken words can bias emotion interpretation (e.g., uttering happy phrases in a flat or sad tone), so STMIL minimizes the mutual information between the speech representation QeQ_e and the transcription representation QtQ_t.

    The mutual information between QtQ_t and QeQ_e is defined as: I(Qt;Qe)=∑qt∈Qt∑qe∈Qep(qt,qe)log⁡p(qt,qe)p(qt)p(qe)I(Q_t; Q_e) = \sum_{q_t \in Q_t} \sum_{q_e \in Q_e} p(q_t, q_e) \log \frac{p(q_t, q_e)}{p(q_t) p(q_e)} where p(qt,qe)p(q_t, q_e) is the joint distribution and p(qt),p(qe)p(q_t), p(q_e) are marginal distributions.

    Because computing I(Qt;Qe)I(Q_t; Q_e) directly is intractable for continuous high-dimensional features, SECap adopts the variational contrastive log-ratio upper bound (vCLUB) to minimize mutual information during training: U(Qt;Qe)=1n2∑i=1n∑j=1n[log⁡q(yi∣xi)q(yj∣xi)]U(Q_t; Q_e) = \frac{1}{n^2} \sum_{i=1}^n \sum_{j=1}^n \left[ \log \frac{q(y_i \mid x_i)}{q(y_j \mid x_i)} \right] where xix_i represents the ii-th sample of transcription representation QtQ_t, yiy_i and yjy_j represent samples of speech representation QeQ_e, and q(y∣x)q(y \mid x) is a neural network estimator modeling the conditional probability of QeQ_e given QtQ_t.

  5. Knowl 5 — Speech-Caption Contrastive Learning (SCCL)

    model/method

    Speech-Caption Contrastive Learning (SCCL) aligns speech representations with emotion caption semantics in SECap's Bridge-Net by minimizing feature distance between speech representations QeQ_e and emotion caption representations QcQ_c.

    The dataset is partitioned into NN distinct emotion categories using human emotion labels. During each training step, KK speech-caption pairs are sampled from each of the NN categories (NKNK total pairs). For each speech embedding ei∈Qee_i \in Q_e:

    • did_i is its paired ground-truth caption embedding (11 sample).
    • pijp_{ij} are caption embeddings from the same emotion category (K−1K - 1 samples).
    • uiju_{ij} are caption embeddings from different emotion categories (NK−KNK - K samples).

    Using cosine similarity S(⋅,⋅)S(\cdot, \cdot), the contrastive loss is defined as: L(Qc;Qe)=∑i=1NK[w1(1−S(ei,di))+w2∑j=1K−1(1−S(ei,pij))+w3∑j=1NK−KReLU(S(ei,uij)−m)]L(Q_c; Q_e) = \sum_{i=1}^{NK} \left[ w_1 (1 - S(e_i, d_i)) + w_2 \sum_{j=1}^{K-1} (1 - S(e_i, p_{ij})) + w_3 \sum_{j=1}^{NK-K} \text{ReLU}(S(e_i, u_{ij}) - m) \right] where w1,w2,w3w_1, w_2, w_3 are weighting hyperparameters controlling each loss component, and mm is a distance margin threshold enforcing separation from dissimilar emotion captions.

  6. Knowl 6 — Two-Stage Training Process of SECap

    model/method

    SECap is trained in two sequential stages:

    1. Stage 1 (Feature Extraction & Disentanglement): The HuBERT audio encoder is frozen. The Q-Former is initialized from pre-trained BERTbase\text{BERT}_{\text{base}} parameters and trained using the combined STMIL and SCCL loss: LT1=wT1×U(Qt;Qe)+wT2×L(Qc;Qe)\mathcal{L}_{T1} = w_{T1} \times U(Q_t; Q_e) + w_{T2} \times L(Q_c; Q_e) where U(Qt;Qe)U(Q_t; Q_e) is the mutual information upper bound loss, L(Qc;Qe)L(Q_c; Q_e) is the contrastive loss, and wT1,wT2w_{T1}, w_{T2} are balance weights. This stage trains 100M parameters out of 517M total parameters.

    2. Stage 2 (LLM Alignment and Generation): The HuBERT audio encoder and LLaMA text decoder parameters are kept frozen. The Q-Former and the linear projection layer (103M trainable parameters out of 7.4B total parameters) are fine-tuned using cross-entropy loss with teacher forcing: LT2=CELoss(C,C^)\mathcal{L}_{T2} = \text{CELoss}(C, \hat{C}) where CC is the ground-truth Chinese caption and C^\hat{C} is the predicted caption. To boost prompt generalization, one of 30 semantically equivalent Chinese instruction prompts (directing the model to describe the speaker's emotion in a single sentence) is randomly selected during each training step.

  7. Knowl 7 — EMOSpeech Dataset and Annotation Protocol

    experimental setup

    The EMOSpeech dataset is a Chinese speech emotion dataset consisting of 41.6 hours of audio sampled at 24 kHz across 30,526 utterances recorded from 7 speakers (5 female, 2 male). The dataset is split into:

    • Training set: 29,326 utterances
    • Validation set: 600 utterances
    • Test set: 600 utterances

    Each audio utterance is annotated with three to five human-written natural language emotion captions, categorical emotion labels, and verbatim speech transcriptions. Annotations are produced using a three-level protocol:

    1. Identification of the coarse emotion using a single categorical word.
    2. Assessment and description of emotion intensity.
    3. Composition of a descriptive sentence combining emotional state, perceived volume, and speech rate.
  8. Knowl 8 — Objective Evaluation of SECap Across Input Modalities

    empirical result

    SECap was objectively evaluated on the EMOSpeech test set against the HTSAT-BART baseline and across different input configurations. Evaluation metrics include sentence similarity metrics SIM1\text{SIM}_1 (MacBERT trained on Chinese STS-B) and SIM2\text{SIM}_2 (Sentence-BERT fine-tuned on Tencent Cloud), alongside automated captioning metrics BLEU1_1, BLEU4_4, METEOR, ROUGEl_l, CIDEr, and SPICE.

    Model ID Input Modality SIM1\text{SIM}_1 SIM2\text{SIM}_2 BLEU1\text{BLEU}_1 BLEU4\text{BLEU}_4 METEOR ROUGEl\text{ROUGE}_l CIDEr SPICE
    Text Audio
    HTSAT-BART #1 — Q-Emb 59.62 53.19 32.74 3.05 14.61 23.64 2.21 2.17
    SECap #2 Raw Trans — 56.49 22.38 0.014 0.00 2.67 4.38 0.00 0.00
    SECap #3 T-Emb — 65.90 62.24 25.97 4.28 15.98 23.37 18.37 2.58
    SECap #4 Raw Trans Q-Emb 69.24 67.50 29.59 5.36 16.99 25.16 28.51 5.77
    SECap #5 T-Emb Q-Emb 69.66 70.02 33.62 7.25 18.44 27.18 33.82 5.96
    SECap #6 — Q-Emb 71.95 70.51 36.08 8.12 19.30 28.49 34.81 6.49

    Key findings:

    1. SECap using speech embeddings alone (#6) outperforms HTSAT-BART (#1) across all objective metrics (SIM1\text{SIM}_1: 71.95 vs 59.62; BLEU4_4: 8.12 vs 3.05; CIDEr: 34.81 vs 2.21).
    2. Using Q-Embedding alone (#6) achieves higher objective scores than combining audio with text representations (#4, #5), because conflicting or redundant semantic information between transcription and audio can complicate LLM feature integration under reference-matching metrics.
  9. Knowl 9 — Subjective Mean Opinion Score Evaluation of Speech Emotion Captions

    empirical result

    A subjective Mean Opinion Score (MOS) study with 15 human evaluators was conducted on 50 randomly selected test utterances. Evaluators scored generated descriptions on a 1-to-5 scale using a three-step protocol: checking whether the sentence describes an emotion, whether the summarized emotion matches the speech audio, and whether the emotional intensity aligns with the audio.

    Reported MOS scores:

    • Human Caption (Ground Truth): 3.85
    • SECap (Q-Embedding + T-Embedding): 3.77
    • SECap (Q-Embedding): 3.69
    • SECap (T-Embedding): 3.59
    • SECap (Raw Trans + Q-Embedding): 3.43
    • Human Label (Categorical): 3.39
    • SER Model Label (Categorical baseline): 3.04
    • HTSAT-BART Baseline: 2.41
    • SECap (Raw Trans): 1.26

    Natural language captions generated by SECap with audio and text embeddings (3.77) perform on par with human-written captions (3.85) and significantly outperform categorical human labels (3.39) and pre-trained SER classification models (3.04).

  10. Knowl 10 — Ablation Study of Model Components in SECap

    empirical result

    An ablation study evaluated the relative contributions of the audio encoder, bridging module, and text decoder to caption quality, measured by SIM1\text{SIM}_1 on the EMOSpeech dataset:

    Audio Encoder Bridge-Net Text Decoder SIM1\text{SIM}_1
    HTSAT Linear BART 59.62±0.2259.62 \pm 0.22
    HuBERT Linear BART 63.95±0.2763.95 \pm 0.27
    HTSAT Linear LLaMA 64.94±0.0664.94 \pm 0.06
    HuBERT Linear LLaMA 68.62±0.3268.62 \pm 0.32
    HuBERT Q-Former LLaMA 71.95±0.04\textbf{71.95} \pm \textbf{0.04}

    Relative gains demonstrated:

    1. Replacing the BART text decoder with LLaMA (keeping HTSAT and Linear) improves SIM1\text{SIM}_1 by +8.92%+8.92\% (from 59.62 to 64.94).
    2. Replacing the HTSAT audio encoder with HuBERT (keeping BART and Linear) improves SIM1\text{SIM}_1 by +7.26%+7.26\% (from 59.62 to 63.95).
    3. Replacing the simple Linear projection layer with Q-Former (with HuBERT and LLaMA) increases SIM1\text{SIM}_1 by +4.85%+4.85\% (from 68.62 to 71.95).
  11. Knowl 11 — Ablation Study of Bridge-Net Training Methods

    empirical result

    An ablation study analyzed the impact of Speech-Transcription Mutual Information Learning (STMIL), Speech-Caption Contrastive Learning (SCCL), and Q-Former freezing during Stage 2 fine-tuning on SIM1\text{SIM}_1 scores:

    Model STMIL SCCL Freeze Stage 2 SIM1\text{SIM}_1
    HTSAT-BART — — — 63.95±0.2763.95 \pm 0.27
    SECap 67.29±0.2267.29 \pm 0.22
    SECap ✓ 68.75±0.1268.75 \pm 0.12
    SECap ✓ 69.40±0.1169.40 \pm 0.11
    SECap ✓ ✓ 71.95±0.04\textbf{71.95} \pm \textbf{0.04}
    SECap ✓ ✓ ✓ 57.92±0.1957.92 \pm 0.19

    Observations:

    1. Adding STMIL alone improves SIM1\text{SIM}_1 by +2.17%+2.17\% (68.75 vs 67.29), showing that removing linguistic content correlation improves emotional representation.
    2. Adding SCCL alone improves SIM1\text{SIM}_1 by +3.14%+3.14\% (69.40 vs 67.29).
    3. Applying both STMIL and SCCL yields a +6.92%+6.92\% increase over the un-pretrained Q-Former (71.95 vs 67.29).
    4. Freezing the Q-Former during Stage 2 fine-tuning and training only the linear projection layer causes a −19.50%-19.50\% drop in SIM1\text{SIM}_1 (57.92 vs 71.95), showing that tuning Q-Former with LLaMA guidance is critical for cross-modal alignment.

Coverage note — None was omitted; all key architectural components, training stages, mathematical objectives, dataset specifications, and empirical ablation findings from the paper's contribution are captured.

References

  1. 1.Anderson, P.; Fernando, B.; Johnson, M.; and Gould, S. 2016. Spice: Semantic propositional image caption evaluation. In Computer Vision–ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11-14, 2016, Proceedings, Part V 14, 382–398. Springer.
  2. 2.Banerjee, S.; and Lavie, A. 2005. METEOR: An automatic metric for MT evaluation with improved correlation with human judgments. In Proceedings of the acl workshop on intrinsic and extrinsic evaluation measures for machine translation and/or summarization, 65–72.
  3. 3.Belghazi, M. I.; Baratin, A.; Rajeshwar, S.; Ozair, S.; Bengio, Y.; Courville, A.; and Hjelm, D. 2018. Mutual information neural estimation. In International conference on machine learning, 531–540. PMLR.
  4. 4.Bilen, C¸ .; Ferroni, G.; Tuveri, F.; Azcarreta, J.; and Krstulovi´c, S. 2020. A framework for the robust evaluation of sound event detection. In ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 61–65. IEEE.
  5. 5.Black, S.; Biderman, S.; Hallahan, E.; Anthony, Q.; Gao, L.; Golding, L.; He, H.; Leahy, C.; McDonell, K.; Phang, J.; et al. 2022. Gpt-neox-20b: An open-source autoregressive language model. arXiv preprint arXiv:2204.06745.
  6. 6.Cer, D.; Diab, M.; Agirre, E.; Lopez-Gazpio, I.; and Specia, L. 2017. SemEval-2017 Task 1: Semantic Textual Similarity Multilingual and Crosslingual Focused Evaluation. In Proceedings of the 11th International Workshop on Semantic Evaluation (SemEval-2017), 1–14. Vancouver, Canada: Association for Computational Linguistics.
  7. 7.Chattopadhyay, S.; Dey, A.; Basak, H.; et al. 2020. Optimizing speech emotion recognition using manta-ray based feature selection. arXiv preprint arXiv:2009.08909.
  8. 8.Chen, K.; Wu, Y.; Wang, Z.; Zhang, X.; Nian, F.; Li, S.; and Shao, X. 2020. Audio Captioning Based on Transformer and Pre-Trained CNN. In DCASE, 21–25.
  9. 9.Cheng, P.; Hao, W.; Dai, S.; Liu, J.; Gan, Z.; and Carin, L. 2020. Club: A contrastive log-ratio upper bound of mutual information. In International conference on machine learning, 1779–1788. PMLR.
  10. 10.Cui, Y.; Che, W.; Liu, T.; Qin, B.; and Yang, Z. 2021. Pre-training with whole word masking for chinese bert. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 29: 3504–3514.
  11. 11.Cui, Y.; Yang, Z.; and Yao, X. 2023. Efficient and effective text encoding for chinese llama and alpaca. arXiv preprint arXiv:2304.08177.
  12. 12.Dahake, P. P.; Shaw, K.; et al. 2016. Speaker dependent speech emotion recognition using MFCC and Support Vector Machine. In 2016 International Conference on Automatic Control and Dynamic Optimization Techniques (ICACDOT), 1080–1084.
  13. 13.Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), 4171–4186.
  14. 14.Du, Z.; Qian, Y.; Liu, X.; Ding, M.; Qiu, J.; Yang, Z.; and Tang, J. 2022. GLM: General Language Model Pretraining with Autoregressive Blank Infilling. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 320–335.
  15. 15.El Ayadi, M.; Kamel, M. S.; et al. 2011. Survey on speech emotion recognition: Features, classification schemes, and databases. Pattern recognition, 44(3): 572–587.
  16. 16.Guzhov, A.; Raue, F.; Hees, J.; and Dengel, A. 2022. Audioclip: Extending clip to image, text and audio. In ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 976–980. IEEE.
  17. 17.Han, Q.; Yuan, W.; Liu, D.; Li, X.; and Yang, Z. 2021. Automated Audio Captioning with Weakly Supervised Pre-Training and Word Selection Methods. In DCASE, 6–10.
  18. 18.Hsu, W.-N.; Bolte, B.; Tsai, Y.-H. H.; Lakhotia, K.; Salakhutdinov, R.; and Mohamed, A. 2021. Hubert: Self-supervised speech representation learning by masked prediction of hidden units. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 29: 3451–3460.
  19. 19.Huang, R.; Li, M.; Yang, D.; Shi, J.; Chang, X.; Ye, Z.; Wu, Y.; Hong, Z.; Huang, J.; Liu, J.; et al. 2023. Audiogpt: Understanding and generating speech, music, sound, and talking head. arXiv preprint arXiv:2304.12995.
  20. 20.Jiang, P.; Fu, H.; Tao, H.; Lei, P.; and Zhao, L. 2019. Parallelized convolutional recurrent neural network with spectral features for speech emotion recognition. IEEE Access, 7: 90368–90377.
  21. 21.Khalil, R. A.; Jones, E.; Babar, M. I.; Jan, T.; Zafar, M. H.; and Alhussain, T. 2019. Speech emotion recognition using deep learning techniques: A review. IEEE Access, 7: 117327–117345.
  22. 22.Koh, A.; Fuzhao, X.; and Siong, C. E. 2022. Automated audio captioning using transfer learning and reconstruction latent space similarity regularization. In ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 7722–7726. IEEE.
  23. 23.Kojima, T.; Gu, S. S.; Reid, M.; Matsuo, Y.; and Iwasawa, Y. 2022. Large language models are zero-shot reasoners. Advances in neural information processing systems, 35: 22199–22213.
  24. 24.Li, J.; Li, D.; Savarese, S.; and Hoi, S. 2023. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. arXiv preprint arXiv:2301.12597.
  25. 25.Lin, C.-Y. 2004. Rouge: A package for automatic evaluation of summaries. In Text summarization branches out, 74–81.
  26. 26.Liu, Z.-T.; Xie, Q.; Wu, M.; Cao, W.-H.; Mei, Y.; and Mao, J.-W. 2018. Speech emotion recognition based on an improved brain emotion learning model. Neurocomputing, 309: 145–156.
  27. 27.Mei, X.; Meng, C.; Liu, H.; Kong, Q.; Ko, T.; Zhao, C.; Plumbley, M. D.; Zou, Y.; and Wang, W. 2023. Wavcaps: A chatgpt-assisted weakly-labelled audio captioning dataset for audio-language multimodal research. arXiv preprint arXiv:2303.17395.
  28. 28.Ming, X. 2022. text2vec: A Tool for Text to Vector.
  29. 29.Mohamed, O.; and Aly, S. A. 2021. Arabic speech emotion recognition employing wav2vec2. 0 and hubert based on baved dataset. arXiv preprint arXiv:2110.04425.
  30. 30.Nwe, T. L.; Foo, S. W.; and De Silva, L. C. 2003. Speech emotion recognition using hidden Markov models. Speech communication, 41(4): 603–623.
  31. 31.OpenAI. 2023. GPT-4 Technical Report. arXiv:2303.08774.
  32. 32.Papineni, K.; Roukos, S.; Ward, T.; and Zhu, W.-J. 2002. Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th annual meeting of the Association for Computational Linguistics, 311–318.
  33. 33.Reimers, N.; and Gurevych, I. 2019. Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks. In Inui, K.; Jiang, J.; Ng, V.; and Wan, X., eds., Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, EMNLP-IJCNLP 2019, Hong Kong, China, November 3-7, 2019, 3980–3990. Association for Computational Linguistics.
  34. 34.Scao, T. L.; Fan, A.; Akiki, C.; Pavlick, E.; Ilic, S.; Hesslow, D.; Castagne, R.; Luccioni, A. S.; Yvon, F.; Gall´e, M.; et al. 2022. Bloom: A 176b-parameter open-access multilingual language model. arXiv preprint arXiv:2211.05100.
  35. 35.Shen, Y.; Song, K.; Tan, X.; Li, D.; Lu, W.; and Zhuang, Y. 2023. Hugginggpt: Solving ai tasks with chatgpt and its friends in huggingface. arXiv preprint arXiv:2303.17580.
  36. 36.Singhal, K.; Azizi, S.; Tu, T.; Mahdavi, S. S.; Wei, J.; Chung, H. W.; Scales, N.; Tanwani, A.; Cole-Lewis, H.; Pfohl, S.; et al. 2023. Large language models encode clinical knowledge. Nature, 1–9.
  37. 37.Sutskever, I.; Vinyals, O.; and Le, Q. V. 2014. Sequence to Sequence Learning with Neural Networks. In Ghahramani, Z.; Welling, M.; Cortes, C.; Lawrence, N. D.; and Weinberger, K. Q., eds., Advances in Neural Information Processing Systems 27: Annual Conference on Neural Information Processing Systems 2014, December 8-13 2014, Montreal, Quebec, Canada, 3104–3112.
  38. 38.Touvron, H.; Lavril, T.; Izacard, G.; Martinet, X.; Lachaux, M.-A.; Lacroix, T.; Roziere, B.; Goyal, N.; Hambro, E.; Azhar, F.; et al. 2023. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971.
  39. 39.Van Den Oord, A.; Vinyals, O.; et al. 2017. Neural discrete representation learning. Advances in neural information processing systems, 30.
  40. 40.Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017. Attention is all you need. Advances in neural information processing systems, 30.
  41. 41.Wang, Q.; and Chan, A. B. 2019. Describing like humans: on diversity in image captioning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 4195–4203.
  42. 42.Wu, Y.; Chen, K.; Zhang, T.; Hui, Y.; Berg-Kirkpatrick, T.; and Dubnov, S. 2023. Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation. In ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 1–5. IEEE.
  43. 43.Xu, X.; Wu, M.; and Yu, K. 2022. A comprehensive survey of automated audio captioning. arXiv preprint arXiv:2205.05357.
  44. 44.Xu, Y.; Huang, Q.; Wang, W.; Foster, P.; Sigtia, S.; Jackson, P. J.; and Plumbley, M. D. 2017. Unsupervised feature learning based on deep models for environmental audio tagging. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 25(6): 1230–1241.
  45. 45.Ye, Z.; Wang, H.; Yang, D.; and Zou, Y. 2021. Improving the Performance of Automated Audio Captioning via Integrating the Acoustic and Semantic Information. In Font, F.; Mesaros, A.; Ellis, D. P. W.; Fonseca, E.; Fuentes, M.; and Elizalde, B., eds., Proceedings of the 6th Workshop on Detection and Classification of Acoustic Scenes and Events 2021 (DCASE 2021), Online, November 15-19, 2021, 40–44.
  46. 46.Zhang, B.; Lv, H.; Guo, P.; Shao, Q.; Yang, C.; Xie, L.; Xu, X.; Bu, H.; Chen, X.; Zeng, C.; et al. 2022. Wenetspeech: A 10000+ hours multi-domain mandarin corpus for speech recognition. In ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 6182–6186. IEEE.
  47. 47.Zhang, H.; Li, X.; and Bing, L. 2023. Video-llama: An instruction-tuned audio-visual language model for video understanding. arXiv preprint arXiv:2306.02858.

Citation

MLA
Xu, Y., et al. “SECap: Speech Emotion Captioning with Large Language Model”. arXiv, 2023, http://arxiv.org/abs/2312.10381v3.
APA
Xu, Y., Chen, H., Yu, J., Huang, Q., Wu, Z., Zhang, S., Li, G., Luo, Y., & Gu, R. (2023). SECap: Speech Emotion Captioning with Large Language Model. arXiv. http://arxiv.org/abs/2312.10381v3
Chicago
Xu, Y., H. Chen, J. Yu, et al. 2023. “SECap: Speech Emotion Captioning with Large Language Model”. arXiv. http://arxiv.org/abs/2312.10381v3.
Harvard
Xu, Y. et al. (2023) “SECap: Speech Emotion Captioning with Large Language Model”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2312.10381v3.
Vancouver
1. Xu Y, Chen H, Yu J, Huang Q, Wu Z, Zhang S, Li G, Luo Y, Gu R (2023) SECap: Speech Emotion Captioning with Large Language Model. arXiv

BibTeX

@article{xu2023secap,
  title = {SECap: Speech Emotion Captioning with Large Language Model},
  author = {Xu, Yaoxun and Chen, Hangting and Yu, Jianwei and Huang, Qiaochu and Wu, Zhiyong and Zhang, Shixiong and Li, Guangzhi and Luo, Yi and Gu, Rongzhi},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2312.10381v3},
  eprint = {2312.10381}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF