keyword
frame-level representations
Frame-level representations are feature vectors or embeddings that capture the visual and spatial content of individual frames within a video sequence. In video analysis and computer vision architectures, these representations are typically generated by processing each frame through an image-based feature extractor, such as a convolutional neural network or a spatial vision transformer. They capture static scene information, object attributes, and appearance cues at discrete points in time. By serving as an intermediate stage between raw frame inputs and video-level modeling, frame-level representations provide a structured sequence of spatial features that subsequent temporal layers or pooling mechanisms can aggregate to understand motion, temporal dynamics, and overall video semantics.
1 item

