keyword
Video Vision Transformer
A Video Vision Transformer is a deep learning architecture that applies transformer self-attention mechanisms directly to video data for computer vision tasks such as video classification and action recognition. Extending the design of standard Vision Transformers from static images to temporal sequences, it divides input video clips into spatio-temporal tokens or patches and processes them through layers of self-attention to capture complex patterns across both space and time. Because video inputs produce long sequences of tokens that increase computational demand, these architectures frequently incorporate factorized attention mechanisms to separate spatial and temporal operations, allowing efficient training and scaling while serving as a pure-transformer alternative to traditional 3D convolutional neural networks.
1 item

