Built independently by an author, for readers. Read the story and support ChapterPal

keyword

hierarchical co-attention model

A hierarchical co-attention model is a multimodal deep learning architecture that jointly coordinates attention across two distinct data modalities, such as text and images, while processing at least one modality across multiple granularities. Commonly applied in visual question answering, this architecture simultaneously determines which visual regions to focus on and which textual elements to prioritize, establishing bidirectional relevance between language and visual features. The hierarchical mechanism structures the textual input across progressive levels of representation, typically spanning word, phrase, and sentence levels, which enables the network to align both fine-grained lexical details and broader semantic context with corresponding visual regions for more accurate reasoning.

1 item