Speech emotion captioning is an artificial intelligence and speech processing task that automatically generates descriptive natural language text to explain the emotional states and nuances expressed in human speech. Unlike traditional speech emotion recognition, which assigns vocal audio to predetermined discrete categories or numerical values, speech emotion captioning produces open-ended, contextualized descriptions of vocal tone, mood, intensity, and subtle expressive shifts. The task typically couples audio encoders with generative language models to translate complex acoustic cues and speech characteristics into coherent, human-readable sentences that provide a richer and more flexible representation of human emotion.