keyword
hierarchical video captioning
Hierarchical video captioning is a computer vision and natural language processing task that involves generating descriptive text for video content across multiple levels of temporal granularity and semantic abstraction. Unlike conventional video captioning approaches that produce a single description for a brief clip, hierarchical captioning models the multi-tiered structure of extended video footage. The process typically produces fine-grained descriptions for short clips containing specific atomic actions, intermediate summaries for longer segments composed of multiple related events, and comprehensive overviews for entire long-form videos. This structured approach enables systems to preserve detailed visual actions while capturing broader narrative context and long-range semantic relationships.
1 item

