keyword
utterance-level representations
Utterance-level representations are compact, fixed-dimensional feature vectors that encapsulate the overall semantic, acoustic, or visual information of an entire spoken phrase, sentence, or segmented conversational turn. Unlike frame-level, word-level, or token-level representations that model fine-grained localized steps in a sequence, utterance-level embeddings summarize the complete duration of an utterance into a single unified profile. In natural language processing, speech processing, and multimodal machine learning, these representations are typically produced by aggregating temporal sequences through pooling mechanisms, attention networks, or neural encoders across text, audio, and visual streams. They serve as essential inputs for downstream tasks that assign a single global prediction to a complete segment of communication, including sentiment analysis, emotion recognition, speaker identification, and dialogue understanding.
1 item

