MeMViT: Memory-Augmented Multiscale Vision Transformer for Efficient Long-Term Video Recognition
Chao-Yuan WuYanghao LiKarttikeya MangalamHaoqi FanBo XiongJitendra MalikChristoph Feichtenhofer
Presents a memory-augmented multiscale vision transformer that processes videos sequentially and caches past representations, extending temporal context length by thirty times with only a four-and-a-half percent computational increase to achieve state-of-the-art long-term video recognition.
Modern video recognition systems perform well on short clips under five seconds but struggle to process long video sequences due to prohibitive computational costs and memory bottlenecks. Standard approaches attempt to capture longer durations by ingesting more frames simultaneously, which leads to exponential increases in computational demand and hardware usage. This limitation hinders practical deployment in continuous, real-time applications such as robotics, augmented reality, and live video analytics.
The article demonstrates an efficient approach for long-term video recognition by introducing MeMViT, a Memory-Augmented Multiscale Vision Transformer. The core objective is to evaluate whether caching and referencing compact representations of past video segments enables a model to maintain extended temporal context without suffering the severe computational overhead of conventional methods.
To achieve this, the authors designed an architecture that processes videos sequentially in an online manner. Rather than processing full videos at once, the system caches internal transformer key and value representations from prior short clips. Current clips attend hierarchically to these cached memories across network layers. To control memory footprint and compute, the approach incorporates a pipelined memory compression mechanism that learns to discard redundant information across time while avoiding the complexity of backpropagation through time. The evaluation benchmarks the architecture across several standard datasets, including AVA for action localization as well as EPIC-Kitchens-100 for action classification and action anticipation.
The findings establish that MeMViT achieves a 30-fold increase in temporal support with only a 4.5% increase in compute, compared to the over 3,000% compute increase required by traditional frame-scaling approaches. The model consistently outperformed existing baselines across all evaluated benchmarks, achieving state-of-the-art results. On the EPIC-Kitchens-100 dataset, it improved action classification accuracy to 46.2% while running approximately three times faster and consuming two to five times less GPU memory than leading mobile video models. In action anticipation, the extended historical context produced significant improvements in predicting upcoming verbs (improving accuracy by 3.5% overall and 3.7% on rare tail actions), outperforming complex multi-modal systems using only standard video pixels.
These results indicate that video recognition models can scale to much longer temporal contexts without requiring specialized multi-model pipelines or proportional increases in hardware infrastructure. By operating on a single backbone with linear computational scaling, the method lowers operational hardware costs, reduces memory footprints, and provides a direct path toward real-time, low-latency video streaming analysis.
For practical application, engineering teams and decision-makers evaluating long-form or streaming video systems should consider adopting online memory-augmented transformer architectures rather than scaling raw input frame counts. Teams can optimize performance and computational trade-offs by applying memory augmentation to a subset of layers (e.g., 50% alternating layers) and incorporating aggressive temporal compression.
The reported findings provide high confidence within the scope of standardized action recognition and egocentric benchmarks. However, the evaluation focuses on sequential video clips within fixed window sizes (up to roughly 70 seconds of receptive field), meaning additional validation is recommended before deploying the architecture to ultra-long temporal horizons spanning hours or to video domains with significant structural differences from the evaluated datasets.
- Paper: Multiscale Vision Transformers, Haoqi Fan et al. (2021). This paper introduces the base Multiscale Vision Transformer (MViT) architecture that MeMViT directly adapts and equips with temporal memory caching.
- Paper: Is Space-Time Attention All You Need for Video Understanding?, Gedas Bertasius et al. (2021). It introduces space-time self-attention schemes for video transformers, establishing the baseline self-attention mechanisms that MeMViT scales to long temporal horizons.
- Paper: ViViT: A Video Vision Transformer, Anurag Arnab et al. (2021). It formulates factorized spatio-temporal attention architectures for video classification, providing essential background on transformer designs for video.
- Paper: TSM: Temporal Shift Module for Efficient Video Understanding, Ji Lin et al. (2018). It demonstrates how caching past temporal features in an online fashion enables efficient long-range context modeling in video recognition.
- Paper: SlowFast Networks for Video Recognition, Christoph Feichtenhofer et al. (2018). It establishes key principles of multi-rate temporal modeling in video recognition that motivate multi-scale and long-range video architectures.
- Paper: Video ReCap: Recursive Captioning of Hour-Long Videos, Md Mohaiminul Islam et al. (2024). This work scales long-range video understanding from seconds-long clips to hour-long videos through hierarchical recursive captioning.
- Paper: HourVideo: 1-Hour Video-Language Understanding, Keshigeyan Chandrasegaran et al. (2024). This benchmark evaluates multimodal models on hour-long egocentric video understanding, pushing long-temporal context modeling beyond short-clip classification.
- Paper: TimeSuite: Improving MLLMs for Long Video Understanding via Grounded Tuning, Xiangyu Zeng 0004 et al. (2025). It advances long-form video reasoning by introducing temporal grounding and compression strategies for multimodal large language models.
- Paper: Efficient Movie Scene Detection using State-Space Transformers, Md Mohaiminul Islam et al. (2023). It explores alternative long-range sequence modeling architectures using state-space transformers to efficiently process full-length movie scenes.
- Paper: Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis, Chaoyou Fu et al. (2025). It presents a comprehensive benchmark evaluating how modern multimodal architectures handle varying temporal durations and long-form video comprehension.
