keyword
audio-visual speech recognition
Audio-visual speech recognition is a multimodal technology that automatically transcribes spoken language into text by combining auditory signals with visual cues, most notably the movement of a speaker's lips and face. By supplementing acoustic data with visual information, which is unaffected by background noise or acoustic distortion, these systems resolve phonetic ambiguities and achieve greater transcription accuracy than traditional audio-only approaches. Modern systems typically use multimodal deep learning architectures to align and fuse features extracted from both audio waveforms and video streams, emulating the human ability to integrate sound and sight for speech comprehension in noisy or complex listening environments.
1 item

