keyword
Temporal Transformer Encoder
A temporal transformer encoder is a neural network component designed to model time-dependent dependencies and dynamic interactions across sequential feature representations. In video understanding and sequence processing architectures that factorize space and time, it operates on sequence-level tokens or frame representations, such as aggregated feature vectors extracted from individual video frames by a preceding spatial encoder. By applying multi-head self-attention specifically across the temporal dimension rather than computing joint attention across all spatial and temporal tokens simultaneously, the temporal transformer encoder captures motion patterns, transitions, and chronological context while substantially reducing computational complexity.
1 item

