MeMViT: Memory-Augmented Multiscale Vision Transformer for Efficient Long-Term Video Recognition

Chao-Yuan WuYanghao LiKarttikeya MangalamHaoqi FanBo XiongJitendra MalikChristoph Feichtenhofer

article2022CVPR284 citations

Presents a memory-augmented multiscale vision transformer that processes videos sequentially and caches past representations, extending temporal context length by thirty times with only a four-and-a-half percent computational increase to achieve state-of-the-art long-term video recognition.

Listen

Modern video recognition systems perform well on short clips under five seconds but struggle to process long video sequences due to prohibitive computational costs and memory bottlenecks. Standard approaches attempt to capture longer durations by ingesting more frames simultaneously, which leads to exponential increases in computational demand and hardware usage. This limitation hinders practical deployment in continuous, real-time applications such as robotics, augmented reality, and live video analytics.

The article demonstrates an efficient approach for long-term video recognition by introducing MeMViT, a Memory-Augmented Multiscale Vision Transformer. The core objective is to evaluate whether caching and referencing compact representations of past video segments enables a model to maintain extended temporal context without suffering the severe computational overhead of conventional methods.

To achieve this, the authors designed an architecture that processes videos sequentially in an online manner. Rather than processing full videos at once, the system caches internal transformer key and value representations from prior short clips. Current clips attend hierarchically to these cached memories across network layers. To control memory footprint and compute, the approach incorporates a pipelined memory compression mechanism that learns to discard redundant information across time while avoiding the complexity of backpropagation through time. The evaluation benchmarks the architecture across several standard datasets, including AVA for action localization as well as EPIC-Kitchens-100 for action classification and action anticipation.

The findings establish that MeMViT achieves a 30-fold increase in temporal support with only a 4.5% increase in compute, compared to the over 3,000% compute increase required by traditional frame-scaling approaches. The model consistently outperformed existing baselines across all evaluated benchmarks, achieving state-of-the-art results. On the EPIC-Kitchens-100 dataset, it improved action classification accuracy to 46.2% while running approximately three times faster and consuming two to five times less GPU memory than leading mobile video models. In action anticipation, the extended historical context produced significant improvements in predicting upcoming verbs (improving accuracy by 3.5% overall and 3.7% on rare tail actions), outperforming complex multi-modal systems using only standard video pixels.

These results indicate that video recognition models can scale to much longer temporal contexts without requiring specialized multi-model pipelines or proportional increases in hardware infrastructure. By operating on a single backbone with linear computational scaling, the method lowers operational hardware costs, reduces memory footprints, and provides a direct path toward real-time, low-latency video streaming analysis.

For practical application, engineering teams and decision-makers evaluating long-form or streaming video systems should consider adopting online memory-augmented transformer architectures rather than scaling raw input frame counts. Teams can optimize performance and computational trade-offs by applying memory augmentation to a subset of layers (e.g., 50% alternating layers) and incorporating aggressive temporal compression.

The reported findings provide high confidence within the scope of standardized action recognition and egocentric benchmarks. However, the evaluation focuses on sequential video clips within fixed window sizes (up to roughly 70 seconds of receptive field), meaning additional validation is recommended before deploying the architecture to ultra-long temporal horizons spanning hours or to video domains with significant structural differences from the evaluated datasets.

arXiv: 2201.08383
  • Paper: Multiscale Vision Transformers, Haoqi Fan et al. (2021). This paper introduces the base Multiscale Vision Transformer (MViT) architecture that MeMViT directly adapts and equips with temporal memory caching.
  • Paper: Is Space-Time Attention All You Need for Video Understanding?, Gedas Bertasius et al. (2021). It introduces space-time self-attention schemes for video transformers, establishing the baseline self-attention mechanisms that MeMViT scales to long temporal horizons.
  • Paper: ViViT: A Video Vision Transformer, Anurag Arnab et al. (2021). It formulates factorized spatio-temporal attention architectures for video classification, providing essential background on transformer designs for video.
  • Paper: TSM: Temporal Shift Module for Efficient Video Understanding, Ji Lin et al. (2018). It demonstrates how caching past temporal features in an online fashion enables efficient long-range context modeling in video recognition.
  • Paper: SlowFast Networks for Video Recognition, Christoph Feichtenhofer et al. (2018). It establishes key principles of multi-rate temporal modeling in video recognition that motivate multi-scale and long-range video architectures.
Cover for MeMViT: Memory-Augmented Multiscale Vision Transformer for Efficient Long-Term Video Recognition

Abstract

While today’s video recognition systems parse snapshots or short clips accurately, they cannot connect the dots and reason across a longer range of time yet. Most existing video architectures can only process <5 seconds of a video without hitting the computation or memory bottlenecks.

In this paper, we propose a new strategy to overcome this challenge. Instead of trying to process more frames at once like most existing methods, we propose to process videos in an online fashion and cache “memory” at each iteration. Through the memory, the model can reference prior context for long-term modeling, with only a marginal cost. Based on this idea, we build MeMViT, a Memory-augmented Multiscale Vision Transformer, that has a temporal support 30× longer than existing models with only 4.5% more compute; traditional methods need >3,000% more compute to do the same. On a wide range of settings, the increased temporal support enabled by MeMViT brings large gains in recognition accuracy consistently. MeMViT obtains state-of-the-art results on the AVA, EPIC-Kitchens-100 action classification, and action anticipation datasets. Code and models will be made publicly available.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Preliminaries
  • 4 MeMViT for Efficient Long-Term Modeling
  • 4.1 Memory Attention and Caching
  • 4.2 Memory Compression
  • 4.3 Implementation Details
  • 5 Experiments
  • 5.1 Scaling Strategies
  • 5.2 Ablation Experiments
  • 5.3 Generalization Analysis
  • 5.4 State-of-the-Art Comparison
  • 6 Conclusion
  • A Appendix
  • A.1 Architecture Specifications
  • A.2 Kinetics Pre-training Details
  • A.3 AVA Experiments
  • A.4 EPIC-Kitchens-100 Experiments
  • A.5 Supplementary Experiments
  • References

Knowls

  1. Knowl 1 — MeMViT Memory Attention Architecture

    model/method

    MeMViT (Memory-augmented Multiscale Vision Transformer) models long video sequences by processing them sequentially as consecutive short spatio-temporal clips X(t)∈RT×H×W×CX^{(t)} \in \mathbb{R}^{T \times H \times W \times C} at time step tt, while caching and referencing internal transformer representations from earlier time steps t′<tt' < t.

    In standard Multiscale Vision Transformers (MViT), pooling is performed after linear projections: Q=PoolQ(XWQ)Q = \text{Pool}_Q(X W_Q), K=PoolK(XWK)K = \text{Pool}_K(X W_K), and V=PoolV(XWV)V = \text{Pool}_V(X W_V). MeMViT swaps the order of pooling and linear projections so that pooling is applied directly to the layer input tensor XX:

    Qˉ=PoolQ(X),Kˉ=PoolK(X),Vˉ=PoolV(X)\bar{Q} = \text{Pool}_Q(X), \quad \bar{K} = \text{Pool}_K(X), \quad \bar{V} = \text{Pool}_V(X)

    Q=QˉWQ,K=KˉWK,V=VˉWVQ = \bar{Q} W_Q, \quad K = \bar{K} W_K, \quad V = \bar{V} W_V

    where PoolQ,PoolK,PoolV\text{Pool}_Q, \text{Pool}_K, \text{Pool}_V are pooling layers downsampling spatiotemporal dimensions, and WQ,WK,WV∈Rd×dW_Q, W_K, W_V \in \mathbb{R}^{d \times d} are learnable projection weights. Performing pooling prior to linear projection reduces the token count passed to the linear projections, lowering computation.

    To reference past context across MM historical iterations without backpropagation through time (BPTT), cached key and value tensors from time steps t−Mt-M to t−1t-1 are concatenated with the current step's pooled keys and values along the token dimension:

    Kˉ(t):=[sg(Kˉ(t−M)),…,sg(Kˉ(t−1)),Kˉ(t)]\bar{K}^{(t)} := \left[ \text{sg}\left(\bar{K}^{(t-M)}\right), \dots, \text{sg}\left(\bar{K}^{(t-1)}\right), \bar{K}^{(t)} \right]

    Vˉ(t):=[sg(Vˉ(t−M)),…,sg(Vˉ(t−1)),Vˉ(t)]\bar{V}^{(t)} := \left[ \text{sg}\left(\bar{V}^{(t-M)}\right), \dots, \text{sg}\left(\bar{V}^{(t-1)}\right), \bar{V}^{(t)} \right]

    where sg(⋅)\text{sg}(\cdot) denotes the stop-gradient operator. The current query Q(t)=Qˉ(t)WQQ^{(t)} = \bar{Q}^{(t)} W_Q then attends to the extended keys K(t)=Kˉ(t)WKK^{(t)} = \bar{K}^{(t)} W_K and values V(t)=Vˉ(t)WVV^{(t)} = \bar{V}^{(t)} W_V via self-attention:

    Z=Softmax(Q(t)(K(t))⊤d)V(t)Z = \text{Softmax}\left(\frac{Q^{(t)} (K^{(t)})^\top}{\sqrt{d}}\right) V^{(t)}

  2. Knowl 2 — Pipelined Memory Compression

    model/method

    To reduce the GPU memory and computational cost of attending over cached key and value tensors from MM historical steps, MeMViT uses learnable compression modules fKf_K and fVf_V (e.g., learnable pooling operators downsampling spatio-temporal dimensions). Naively optimizing fKf_K and fVf_V across all MM past steps at every iteration would require retaining uncompressed tensors for all historical steps in GPU memory during training.

    MeMViT resolves this through pipelined memory compression. In this formulation, only the memory Kˉ(t−1)\bar{K}^{(t-1)} from the immediately preceding step is cached uncompressed and used to train fKf_K at the current step tt. Memories from earlier steps t′∈[t−M,t−2]t' \in [t-M, t-2] are cached in their already-compressed states K^(t′)\hat{K}^{(t')} from earlier iterations. The augmented key representation at iteration tt is:

    Kˉ(t):=[K^(t−M),…,K^(t−2),fK(sg(Kˉ(t−1))),Kˉ(t)]\bar{K}^{(t)} := \left[ \hat{K}^{(t-M)}, \dots, \hat{K}^{(t-2)}, f_K\left(\text{sg}\left(\bar{K}^{(t-1)}\right)\right), \bar{K}^{(t)} \right]

    where sg(⋅)\text{sg}(\cdot) denotes the stop-gradient operator, Kˉ(t)\bar{K}^{(t)} is the pooled key tensor of the current clip, and K^(t′)=sg(fK(Kˉ(t′)))\hat{K}^{(t')} = \text{sg}(f_K(\bar{K}^{(t')})). The exact same pipelined caching and compression mechanism is applied to values Vˉ(t)\bar{V}^{(t)} using compression module fVf_V.

    Because compression is performed on only one step (Kˉ(t−1)\bar{K}^{(t-1)}) per forward pass, the computational cost of training the compression module is O(1)O(1) with respect to memory length MM, while all older cached tokens maintain an aggressive reduction in attention compute and caching footprint.

  3. Knowl 3 — MeMViT Attention Algorithm with Pipelined Memory Caching

    algorithm

    The forward pass of a MeMViT attention layer with pipelined memory caching and compression over sequential video clips is defined by the following procedure:

    Input: Input tensor xx for the current clip at time step tt
    Input: Pooling layers pool_q,pool_k,pool_v\text{pool\_q}, \text{pool\_k}, \text{pool\_v}
    Input: Linear projection layers lin_q,lin_k,lin_v\text{lin\_q}, \text{lin\_k}, \text{lin\_v}
    Input: Compression modules fk,fvf_k, f_v
    Input: Memory queues mk,mvm_k, m_v storing past keys and values
    Input: Maximum memory length max_len\text{max\_len} (number of historical clips MM)
    Output: Attended output tensor zz
    q←pool_q(x)q \leftarrow \text{pool\_q}(x)
    k←pool_k(x)k \leftarrow \text{pool\_k}(x)
    v←pool_v(x)v \leftarrow \text{pool\_v}(x)
    if length of mk>0m_k > 0:
        cmk←fk(mk[−1])cm_k \leftarrow f_k(m_k[-1])
        cmv←fv(mv[−1])cm_v \leftarrow f_v(m_v[-1])
        past_k←mk[0:−1]+[cmk]past\_k \leftarrow m_k[0 : -1] + [cm_k]
        past_v←mv[0:−1]+[cmv]past\_v \leftarrow m_v[0 : -1] + [cm_v]
        mk[−1]←detach(cmk)m_k[-1] \leftarrow \text{detach}(cm_k)
        mv[−1]←detach(cmv)m_v[-1] \leftarrow \text{detach}(cm_v)
    else:
        past_k←[]past\_k \leftarrow []
        past_v←[]past\_v \leftarrow []
    aug_k←concatenate(past_k+[k],dimension=token)aug\_k \leftarrow \text{concatenate}(past\_k + [k], \text{dimension}=\text{token})
    aug_v←concatenate(past_v+[v],dimension=token)aug\_v \leftarrow \text{concatenate}(past\_v + [v], \text{dimension}=\text{token})
    Q←lin_q(q)Q \leftarrow \text{lin\_q}(q)
    K←lin_k(aug_k)K \leftarrow \text{lin\_k}(aug\_k)
    V←lin_v(aug_v)V \leftarrow \text{lin\_v}(aug\_v)
    z←Softmax(QK⊤/d)Vz \leftarrow \text{Softmax}(Q K^\top / \sqrt{d}) V
    Append detach(k)\text{detach}(k) to mkm_k
    Append detach(v)\text{detach}(v) to mvm_v
    while length of mk>max_len+1m_k > \text{max\_len} + 1:
        Remove oldest element from mkm_k
        Remove oldest element from mvm_v
    return zz

    At any iteration, mk[−1]m_k[-1] holds the uncompressed keys from step t−1t-1, while elements mk[0:−2]m_k[0 : -2] hold detached compressed keys from steps t−Mt-M through t−2t-2.

  4. Knowl 4 — Receptive Field and Complexity Scaling in MeMViT

    theoretical result

    Traditional long-term video architectures increase temporal support by scaling the number of input frames TT processed jointly in a single forward pass, which scales attention compute and activation memory quadratically as O(T2)O(T^2).

    MeMViT scales temporal support linearly and hierarchically:

    1. Complexity Growth: By caching key and value states with stop-gradient operators, prior clip representations are not recomputed, and feed-forward networks (MLPs) only process current clip tokens. Only the attention layer scales with the cached token count, resulting in computational and memory scaling that is O(M)O(M) with respect to the memory span MM.
    2. Hierarchical Receptive Field Growth: In a multi-layer transformer with LL layers, each layer ll at step tt attends to cached states from layer l−1l-1 of preceding clips t′<tt' < t. Because layer l−1l-1 at step t−1t-1 already incorporated information from earlier steps, the temporal receptive field grows hierarchically across network depth LL and memory length MM. For input clips of 2.25 seconds, a memory length of M=2M=2 provides a 16×16\times larger receptive field (36 seconds), and M=4M=4 provides a 32×32\times larger receptive field (70.4 seconds).
  5. Knowl 5 — Relative Positional Embeddings and Video Boundary Resetting in MeMViT

    model/method

    MeMViT incorporates two specific structural adaptations for sequential streaming video processing:

    1. Relative Positional Embeddings: Standard vision transformers utilize absolute positional embeddings defined over tokens of a single input clip. Because consecutive clips tt and t′t' share identical internal coordinates, absolute positional embeddings cannot represent cross-clip temporal offsets. MeMViT employs relative positional embeddings across space and time, enabling queries at time step tt to explicitly distinguish the temporal distance of memory tokens cached from different preceding time steps t′<tt' < t.

    2. Video Boundary Resetting: When processing continuous streaming video or concatenated datasets sequentially during training and inference, consecutive iterations may cross from the end of one video to the beginning of another. Whenever an iteration crosses a video boundary, the cached memory buffers for keys and values are masked to zero (or emptied), preventing memory representations from leaking across distinct video sequences.

  6. Knowl 6 — Computational and Memory Scaling Comparison Between MeMViT and Input Frame Scaling

    empirical result

    Comparing MeMViT with the standard baseline scaling method (which increases temporal context by increasing input frame count TT) on the AVA action localization dataset demonstrates significant compute and memory advantages:

    • Temporal Support and Compute Trade-off: MeMViT expands temporal support by 30×30\times (from ∼2.25\sim 2.25 seconds to >40>40 seconds) with a 4.5%4.5\% increase in FLOPs (from 57.4 GFLOPs to 60.0 GFLOPs). In contrast, scaling input frame count directly requires >3000%>3000\% more compute to achieve comparable temporal durations.
    • GPU Memory and Latency: As temporal support increases from 2 to 60 seconds, input frame scaling causes training GPU memory to rise from 6 GB to over 12 GB, inference GPU memory to rise from 3 GB to over 5 GB, training iteration time to increase from 0.8s to >2.0s, and inference iteration time from 0.1s to 0.25s. Under MeMViT, training memory remains bounded between 6 and 7 GB, test memory remains near 3.2 GB, training iteration time stays near 0.9s, and test iteration time stays near 0.12s across the full 60-second support range.
    • Accuracy at Fixed Compute: At an equal computational budget (∼59\sim 59 GFLOPs), MeMViT achieves 29.3% mAP on AVA, whereas the baseline short-term model without memory achieves 27.0% mAP.
  7. Knowl 7 — Ablation of Architectural Design Choices on AVA Dataset

    data/table

    Ablation experiments on the AVA v2.2 spatio-temporal action localization benchmark using an MViTv2-B backbone pre-trained on Kinetics-400 analyze per-layer memory length MM, memory compression downsampling factors (T×H×WT \times H \times W), and memory augmentation layer placement:

    Ablation Factor Setting GFLOPs mAP (%)
    Memory Length (MM) w/o mem (1×1\times Receptive Field) 57.4 27.0
    M=1M=1 (8×8\times Receptive Field) 58.1 28.7
    M=2M=2 (16×16\times Receptive Field, default) 58.7 29.3
    M=3M=3 (24×24\times Receptive Field) 59.3 29.2
    M=4M=4 (32×32\times Receptive Field) 60.0 28.8
    Compression Factor (T×H×WT \times H \times W) none (uncompressed) 73.0 28.9
    1×2×21 \times 2 \times 2 62.3 29.0
    2×1×12 \times 1 \times 1 65.3 29.1
    2×2×22 \times 2 \times 2 59.9 29.0
    2×4×42 \times 4 \times 4 58.2 28.3
    4×2×24 \times 2 \times 2 (default) 58.7 29.3
    4×4×44 \times 4 \times 4 57.8 28.6
    Augmented Attention Layers all (100%) 60.2 29.1
    75% (uniform) 59.5 29.1
    50% (uniform, default) 58.7 29.3
    25% (uniform) 58.1 28.7
    early (Stage 1 2) 58.4 28.6
    middle (Stage 3) 58.8 28.7
    late (Stage 4) 57.8 29.1

    Key empirical findings:

    • Memory Length: Incorporating memory yields 1.7–2.3% mAP gains over the short-term baseline, with M=2M=2 (16× receptive field, 36s duration) performing best at 29.3% mAP.
    • Compression Factor: Downsampling memory by 4×2×24 \times 2 \times 2 outperforms the uncompressed setting (29.3% vs. 28.9% mAP) while reducing GFLOPs from 73.0 to 58.7, indicating that compression filters irrelevant temporal-spatial noise.
    • Layer Selection: Augmenting 50% of the transformer layers uniformly (alternating standard self-attention and memory-augmented attention) achieves better accuracy than augmenting all layers (29.3% vs. 29.1% mAP) while reducing computation.
  8. Knowl 8 — Spatio-Temporal Action Localization on AVA v2.2 Benchmark

    data/table

    Evaluations on the AVA v2.2 spatio-temporal action localization benchmark compare MeMViT with prior 3D-CNN and transformer architectures across different Kinetics pre-training datasets (Kinetics-400, Kinetics-600, Kinetics-700):

    Model Pre-train Center mAP (%) Full mAP (%) GFLOPs Param (M)
    SlowFast 8×\times8, R101 K400 23.8 - 137.7 53.0
    MViTv1-B, 64×\times3 K400 27.3 - 454.7 36.4
    MViTv2-16, 16×\times4 K400 26.2 27.0 57.4 34.5
    MeMViT-16, 16×\times4 K400 28.5 29.3 58.7 35.4
    SlowFast 16×\times8 R101+NL K600 27.5 - 296.3 59.2
    Object Transformer K600 31.0 - 243.8 86.2
    ACAR 8×\times8, R101-NL K600 - 31.4 ≥293.2\ge 293.2 ≥118.4\ge 118.4
    MViTv2-24, 32×\times3 K600 29.4 30.1 204.4 51.3
    MeMViT-24, 32×\times3 K600 31.5 32.3 211.7 52.6
    MeMViT-24, 32×\times3, ↑3122\uparrow 312^2 K600 32.8 33.6 620.0 52.6
    AIA K700 32.3 - - -
    ACAR R101 K700 - 33.3 ≥212.0\ge 212.0 ≥107.4\ge 107.4
    MViTv2-24, 32×\times3 K700 31.8 32.5 204.4 51.3
    MeMViT-24, 32×\times3 K700 33.5 34.4 211.7 52.6
    MeMViT-24, 32×\times3, ↑3122\uparrow 312^2 K700 34.4 35.4 620.0 52.6

    Under identical backbone architectures, MeMViT consistently outperforms the short-term MViTv2 baseline by +2.3%+2.3\% mAP on K400 (29.3% vs. 27.0%), +2.2%+2.2\% mAP on K600 (32.3% vs. 30.1%), and +1.9%+1.9\% mAP on K700 (34.4% vs. 32.5%). With higher testing resolution (3122312^2), MeMViT-24 reaches 35.4% mAP, outperforming long-term feature-bank methods like ACAR without requiring dual backbones or separate feature extraction stages.

  9. Knowl 9 — EPIC-Kitchens-100 Egocentric Action Classification Benchmark

    data/table

    Evaluations on EPIC-Kitchens-100 action classification assess top-1 accuracy (%) for Action, Verb, and Noun categories:

    Model Pre-train Action (%) Verb (%) Noun (%) Runtime (s) Mem (GB) FLOPs (G) Param (M)
    TSN IN1K 33.2 60.2 46.0 - - - -
    TSM IN1K 38.3 67.9 49.0 - - - -
    SlowFast K400 38.5 65.6 50.0 - - - -
    ViViT-L/16×\times2 IN21K 44.0 66.4 56.8 - - 3410 100
    MFormer-HR IN21K+K400 44.5 67.0 58.5 - - 959 382
    MoViNet-A5 N/A 44.5 69.1 55.1 0.49 8.3 74.9 15.7
    MeMViT, 16×\times4 K400 46.2 70.6 58.5 0.16 1.7 58.7 35.4
    MoViNet-A6 N/A 47.7 72.2 57.3 0.85 8.3 117.0 31.4
    MeMViT, 32×\times3 K600 48.4 71.4 60.3 0.35 3.9 211.7 52.6

    Compared to its short-term counterpart (MViTv2 at 44.6% Action, 69.7% Verb, 56.1% Noun), MeMViT-16 achieves a +1.6%+1.6\% gain on Action and a +2.4%+2.4\% gain on Noun (46.2% Action, 58.5% Noun), indicating that past memory helps disambiguate occluded and blurred objects in egocentric video. MeMViT-24 achieves 48.4% Action accuracy. Additionally, MeMViT-16 operates 3×3\times faster (0.16s vs. 0.85s runtime per clip) with 5×5\times lower GPU memory (1.7 GB vs. 8.3 GB) than MoViNet-A6 during sequential single-pass evaluation.

  10. Knowl 10 — EPIC-Kitchens-100 Egocentric Action Anticipation Benchmark

    data/table

    In egocentric action anticipation on EPIC-Kitchens-100, models observe video up to 1 second before action onset and predict upcoming actions. Models are evaluated using class-mean recall@5 (%) on Overall, Unseen, and Tail splits across Action, Verb, and Noun:

    Model Extra Data / Annotations Overall Unseen Tail
    Action Verb Noun Action Verb Noun Action Verb Noun
    TempAgg IN1K + Boxes + Flow 14.7 23.2 31.4 14.5 28.0 26.2 11.8 14.5 22.5
    RULSTM IN1K + Boxes + Flow 14.0 27.8 30.8 14.2 28.8 27.2 11.1 19.8 22.0
    AVT+ IN21K + Boxes + Multi-loss 15.9 28.2 32.0 11.9 29.5 23.9 14.1 21.1 25.8
    AVT (RGB only) IN21K 14.9 30.2 31.7 - - - - - -
    MeMViT, 16×\times4 K400 15.1 32.8 33.2 9.8 27.5 21.7 13.2 26.3 27.4
    MeMViT, 32×\times3 K700 17.7 32.2 37.0 15.2 28.6 27.4 15.5 25.3 31.0

    Key results:

    • Comparison to Multi-modal Methods: MeMViT-24 (RGB-only trained with standard cross-entropy loss) outperforms AVT+ (which utilizes ImageNet-21K pre-training, object bounding box features, and auxiliary loss functions) by +1.8%+1.8\% on Overall Action (17.7% vs. 15.9%), +4.0%+4.0\% on Overall Verb (32.2% vs. 28.2%), and +5.0%+5.0\% on Overall Noun (37.0% vs. 32.0%).
    • Impact on Verb Anticipation: Compared to the short-term MViTv2 baseline (14.6% Action, 29.3% Verb, 31.8% Noun), MeMViT-16 improves Overall Verb recall by +3.5%+3.5\% (32.8% vs. 29.3%) and Tail Verb recall by +3.7%+3.7\% (26.3% vs. 22.6%), demonstrating that long-term historical context is critical for forecasting rapidly changing actions.

Coverage note — No substantial contributed material was omitted.

References

  1. 1.Sami Abu-El-Haija, Nisarg Kothari, Joonseok Lee, Paul Natsev, George Toderici, Balakrishnan Varadarajan, and Sudheendra Vijayanarasimhan. Youtube-8m: A large-scale video classification benchmark. arXiv:1609.08675, 2016. 2
  2. 2.Anurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun, Mario Lućić, and Cordelia Schmid. ViViT: A video vision transformer. In Proc. ICCV, 2021. 2, 8
  3. 3.Gedas Bertasius, Heng Wang, and Lorenzo Torresani. Is space-time attention all you need for video understanding? In Proc. ICCV, 2021. 2
  4. 4.Joao Carreira, Eric Noland, Andras Banki-Horvath, Chloe Hillier, and Andrew Zisserman. A short note about Kinetics-600. arXiv:1808.01340, 2018. 2, 7
  5. 5.Joao Carreira, Eric Noland, Chloe Hillier, and Andrew Zisserman. A short note on the kinetics-700 human action dataset. arXiv preprint arXiv:1907.06987, 2019. 7, 8
  6. 6.Joao Carreira, Viorica Patraucean, Laurent Mazare, Andrew Zisserman, and Simon Osindero. Massively parallel video networks. In Proc. ECCV, 2018. 2
  7. 7.Shoufa Chen, Peize Sun, Enze Xie, Chongjian Ge, Jiannan Wu, Lan Ma, Jiajun Shen, and Ping Luo. Watch only once: An end-to-end video action detection framework. In Proc. ICCV, 2021. 7
  8. 8.Yihong Chen, Yue Cao, Han Hu, and Liwei Wang. Memory enhanced global-local aggregation for video object detection. In Proc. CVPR, 2020. 2
  9. 9.Zhengsu Chen, Lingxi Xie, Jianwei Niu, Xuefeng Liu, Longhui Wei, and Qi Tian. Visformer: The vision-friendly transformer. In Proc. ICCV, 2021. 2
  10. 10.Changmao Cheng, Chi Zhang, Yichen Wei, and Yu-Gang Jiang. Sparse temporal causal convolution for efficient action modeling. In ACM MM, 2019. 2
  11. 11.Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc V Le, and Ruslan Salakhutdinov. Transformer-xl: Attentive language models beyond a fixed-length context. In ACL, 2019. 2
  12. 12.Navneet Dalal and Bill Triggs. Histograms of oriented gradients for human detection. In Proc. CVPR, 2005. 2
  13. 13.Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Sanja Fidler, Antonino Furnari, Evangelos Kazakos, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, and Michael Wray. Scaling egocentric vision: The epic-kitchens dataset. In ECCV, 2018. 2, 7, 8
  14. 14.Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Sanja Fidler, Antonino Furnari, Evangelos Kazakos, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, and Michael Wray. The epic-kitchens dataset: Collection, challenges and baselines. PAMI, 2021. 2, 7, 8
  15. 15.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. ImageNet: A large-scale hierarchical image database. In Proc. CVPR, 2009. 8
  16. 16.Piotr Dollar, Vincent Rabaud, Garrison Cottrell, and Serge Belongie. Behavior recognition via sparse spatio-temporal features. In International Workshop on Visual Surveillance and Performance Evaluation of Tracking and Surveillance, 2005. 2
  17. 17.Jeff Donahue, Lisa Anne Hendricks, Sergio Guadarrama, Marcus Rohrbach, Subhashini Venugopalan, Kate Saenko, and Trevor Darrell. Long-term recurrent convolutional networks for visual recognition and description. In Proc. CVPR, 2015. 2
  18. 18.Xiaoyi Dong, Jianmin Bao, Dongdong Chen, Weiming Zhang, Nenghai Yu, Lu Yuan, Dong Chen, and Baining Guo. Cswin transformer: A general vision transformer backbone with cross-shaped windows. arXiv preprint arXiv:2107.00652, 2021. 2
  19. 19.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. In Proc. ICLR, 2021. 2
  20. 20.Alexei A Efros, Alexander C Berg, Greg Mori, and Jitendra Malik. Recognizing action at a distance. In Proc. ICCV, 2003. 2
  21. 21.Haoqi Fan, Yanghao Li, Bo Xiong, Wan-Yen Lo, and Christoph Feichtenhofer. PySlowFast. https://github.com/facebookresearch/slowfast, 2020. 5
  22. 22.Haoqi Fan, Bo Xiong, Karttikeya Mangalam, Yanghao Li, Zhicheng Yan, Jitendra Malik, and Christoph Feichtenhofer. Multiscale vision transformers. In Proc. ICCV, 2021. 2, 3, 5, 6, 7
  23. 23.Christoph Feichtenhofer. X3D: Expanding architectures for efficient video recognition. In Proc. CVPR, 2020. 2, 5, 7
  24. 24.Christoph Feichtenhofer, Haoqi Fan, Jitendra Malik, and Kaiming He. SlowFast networks for video recognition. In Proc. ICCV, 2019. 2, 3, 5, 7, 8
  25. 25.Antonino Furnari, Sebastiano Battiato, and Giovanni Maria Farinella. Leveraging uncertainty to rethink loss functions and evaluation measures for egocentric action anticipation. In ECCV Workshops, 2018. 7, 8
  26. 26.Antonino Furnari and Giovanni Farinella. Rolling-unrolling lstms for action anticipation from first-person video. PAMI, 2020. 8
  27. 27.Rohit Girdhar, Joao Carreira, Carl Doersch, and Andrew Zisserman. Video action transformer network. In Proc. CVPR, 2019. 2
  28. 28.Rohit Girdhar and Kristen Grauman. Anticipative Video Transformer. In Proc. ICCV, 2021. 8
  29. 29.Rohit Girdhar, Deva Ramanan, Abhinav Gupta, Josef Sivic, and Bryan Russell. ActionVLAD: Learning spatio-temporal aggregation for action classification. In Proc. CVPR, 2017. 2
  30. 30.Benjamin Graham, Alaaeldin El-Nouby, Hugo Touvron, Pierre Stock, Armand Joulin, Herve Jegou, and Matthijs Douze. LeViT: A vision transformer in ConvNet’s clothing for faster inference. In Proc. ICCV, 2021. 2
  31. 31.Chunhui Gu, Chen Sun, David A. Ross, Carl Vondrick, Caroline Pantofaru, Yeqing Li, Sudheendra Vijayanarasimhan, George Toderici, Susanna Ricco, Rahul Sukthankar, Cordelia Schmid, and Jitendra Malik. AVA: A video dataset of spatio-temporally localized atomic visual actions. In Proc. CVPR, 2018. 2, 5, 6, 7
  32. 32.Noureldien Hussein, Efstratios Gavves, and Arnold WM Smeulders. Timeception for complex action recognition. In Proc. CVPR, 2019. 2
  33. 33.Boyuan Jiang, MengMeng Wang, Weihao Gan, Wei Wu, and Junjie Yan. STM: Spatiotemporal and motion encoding for action recognition. In Proc. CVPR, 2019. 2
  34. 34.Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, et al. The kinetics human action video dataset. arXiv:1705.06950, 2017. 5, 6, 7, 8
  35. 35.Alexander Klaser, Marcin Marszałek, and Cordelia Schmid. A spatio-temporal descriptor based on 3d-gradients. In Proc. BMVC., 2008. 2
  36. 36.Dan Kondratyuk, Liangzhe Yuan, Yandong Li, Li Zhang, Mingxing Tan, Matthew Brown, and Boqing Gong. MoViNets: Mobile video networks for efficient video recognition. In Proc. CVPR, 2021. 2, 8
  37. 37.Bruno Korbar, Du Tran, and Lorenzo Torresani. Scsampler: Sampling salient clips from video for efficient action recognition. In Proc. ICCV, 2019. 2
  38. 38.Ivan Laptev, Marcin Marszalek, Cordelia Schmid, and Benjamin Rozenfeld. Learning realistic human actions from movies. In Proc. CVPR, 2008. 2
  39. 39.Sangmin Lee, Hak Gu Kim, Dae Hwi Choi, Hyung-Il Kim, and Yong Man Ro. Video prediction recalling long-term motion context via memory alignment learning. In Proc. CVPR, 2021. 2
  40. 40.Sangho Lee, Jinyoung Sung, Youngjae Yu, and Gunhee Kim. A memory network approach for story-based temporal summarization of 360 videos. In Proc. CVPR, 2018. 2
  41. 41.Dong Li, Zhaofan Qiu, Qi Dai, Ting Yao, and Tao Mei. Recurrent tubelet proposal and recognition networks for action detection. In Proc. ECCV, 2018. 2
  42. 42.Yanghao Li, Tushar Nagarajan, Bo Xiong, and Kristen Grauman. Ego-exo: Transferring visual representations from third-person to first-person videos. In Proc. CVPR, 2021. 8
  43. 43.Yanghao Li, Chao-Yuan Wu, Haoqi Fan, Karttikeya Mangalam, Bo Xiong, Jitendra Malik, and Christoph Feichtenhofer. Improved multiscale vision transformers for classification and detection. arXiv preprint arXiv:2112.01526, 2021. 2, 5, 7
  44. 44.Zhenyang Li, Kirill Gavrilyuk, Efstratios Gavves, Mihir Jain, and Cees GM Snoek. VideoLSTM convolves, attends and flows for action recognition. Computer Vision and Image Understanding, 2018. 2
  45. 45.Ji Lin, Chuang Gan, and Song Han. TSM: Temporal shift module for efficient video understanding. In Proc. ICCV, 2019. 2, 8
  46. 46.Mason Liu and Menglong Zhu. Mobile video object detection with temporally-aware feature maps. In Proc. CVPR, 2018. 2
  47. 47.Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. In Proc. CVPR, 2022. 2
  48. 48.Daniel Neimark, Omri Bar, Maya Zohar, and Dotan Asselmann. Video transformer network. arXiv preprint arXiv:2102.00719, 2021. 2
  49. 49.Joe Yue-Hei Ng, Matthew Hausknecht, Sudheendra Vijayanarasimhan, Oriol Vinyals, Rajat Monga, and George Toderici. Beyond short snippets: Deep networks for video classification. In Proc. CVPR, 2015. 2
  50. 50.Junting Pan, Siyu Chen, Mike Zheng Shou, Yu Liu, Jing Shao, and Hongsheng Li. Actor-context-actor relation network for spatio-temporal action localization. In Proc. CVPR, 2021. 2, 7, 8
  51. 51.Mandela Patrick, Dylan Campbell, Yuki M Asano, Ishan Misra Florian Metze, Christoph Feichtenhofer, Andrea Vedaldi, Jo Henriques, et al. Keeping your eye on the ball: Trajectory attention in video transformers. In NeurIPS, 2021. 2, 8
  52. 52.Xiaojiang Peng, Changqing Zou, Yu Qiao, and Qiang Peng. Action recognition with stacked fisher vectors. In Proc. ECCV, 2014. 2
  53. 53.Zhaofan Qiu, Ting Yao, and Tao Mei. Learning spatio-temporal representation with pseudo-3d residual networks. In Proc. ICCV, 2017. 2
  54. 54.Ilija Radosavovic, Raj Prateek Kosaraju, Ross Girshick, Kaiming He, and Piotr Dollar. Designing network design spaces. In Proc. CVPR, 2020. 8
  55. 55.Jack W Rae, Anna Potapenko, Siddhant M Jayakumar, and Timothy P Lillicrap. Compressive transformers for long-range sequence modelling. In ICLR, 2019. 2
  56. 56.Jack W Rae and Ali Razavi. Do transformers need deep long-range memory. In ACL, 2020. 2, 6
  57. 57.Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. Faster R-CNN: Towards real-time object detection with region proposal networks. In NIPS, 2015. 2
  58. 58.Fadime Sener, Dibyadip Chatterjee, and Angela Yao. Technical report: Temporal aggregate representations. arXiv preprint arXiv:2106.03152, 2021. 8
  59. 59.Sainbayar Sukhbaatar, Edouard Grave, Piotr Bojanowski, and Armand Joulin. Adaptive attention span in transformers. In ACL, 2019. 2
  60. 60.Sainbayar Sukhbaatar, Da Ju, Spencer Poff, Stephen Roller, Arthur Szlam, Jason Weston, and Angela Fan. Not all memories are created equal: Learning to forget by expiring. In ICML, 2021. 2
  61. 61.Lin Sun, Kui Jia, Kevin Chen, Dit-Yan Yeung, Bertram E Shi, and Silvio Savarese. Lattice long short-term memory for human action recognition. In Proc. ICCV, 2017. 2
  62. 62.Jiajun Tang, Jin Xia, Xinzhi Mu, Bo Pang, and Cewu Lu. Asynchronous interaction aggregation for action detection. In Proc. ECCV, 2020. 7
  63. 63.Graham W Taylor, Rob Fergus, Yann LeCun, and Christoph Bregler. Convolutional learning of spatio-temporal features. In Proc. ECCV, 2010. 2
  64. 64.Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Herve Jegou. Training data-efficient image transformers and distillation through attention. In Proc. ICML, 2021. 2
  65. 65.Hugo Touvron, Matthieu Cord, Alexandre Sablayrolles, Gabriel Synnaeve, and Hervé Jegou. Going deeper with image transformers. In Proc. ICCV, 2021. 2
  66. 66.Du Tran, Lubomir Bourdev, Rob Fergus, Lorenzo Torresani, and Manohar Paluri. Learning spatiotemporal features with 3D convolutional networks. In Proc. ICCV, 2015. 2
  67. 67.Du Tran, Heng Wang, Lorenzo Torresani, and Matt Feiszli. Video classification with channel-separated convolutional networks. In Proc. ICCV, 2019. 2
  68. 68.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. In NeurIPS, 2017. 2
  69. 69.Heng Wang, Alexander Klaser, Cordelia Schmid, and Cheng-Lin Liu. Dense trajectories and motion boundary descriptors for action recognition. IJCV, 2013. 2
  70. 70.Heng Wang and Cordelia Schmid. Action recognition with improved trajectories. In Proc. ICCV, 2013. 2
  71. 71.Heng Wang, Muhammad Muneeb Ullah, Alexander Klaser, Ivan Laptev, and Cordelia Schmid. Evaluation of local spatio-temporal features for action recognition. In BMVC, 2009. 2
  72. 72.Limin Wang, Yu Qiao, and Xiaoou Tang. Action recognition with trajectory-pooled deep-convolutional descriptors. In Proc. CVPR, 2015. 2
  73. 73.Limin Wang, Yuanjun Xiong, Zhe Wang, Yu Qiao, Dahua Lin, Xiaoou Tang, and Luc Val Gool. Temporal segment networks: Towards good practices for deep action recognition. In Proc. ECCV, 2016. 2, 8
  74. 74.Wenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan, Kaitao Song, Ding Liang, Tong Lu, Ping Luo, and Ling Shao. Pyramid vision transformer: A versatile backbone for dense prediction without convolutions. In Proc. ICCV, 2021. 2
  75. 75.Xiaolong Wang, Ross Girshick, Abhinav Gupta, and Kaiming He. Non-local neural networks. In Proc. CVPR, 2018. 2, 3, 5
  76. 76.Xiaohan Wang, Linchao Zhu, Heng Wang, and Yi Yang. Interactive prototype learning for egocentric action recognition. In Proc. ICCV, 2021. 8
  77. 77.Chao-Yuan Wu, Christoph Feichtenhofer, Haoqi Fan, Kaiming He, Philipp Krähenbühl, and Ross Girshick. Long-term feature banks for detailed video understanding. In Proc. CVPR, 2019. 2, 5
  78. 78.Chao-Yuan Wu and Philipp Krahenbuhl. Towards long-form video understanding. In Proc. CVPR, 2021. 2, 7
  79. 79.Chao-Yuan Wu, Manzil Zaheer, Hexiang Hu, R Manmatha, Alexander J Smola, and Philipp Krähenbühl. Compressed video action recognition. In Proc. CVPR, 2018. 2
  80. 80.Saining Xie, Chen Sun, Jonathan Huang, Zhuowen Tu, and Kevin Murphy. Rethinking spatiotemporal feature learning for video understanding. arXiv:1712.04851, 2017. 2
  81. 81.Li Yuan, Yunpeng Chen, Tao Wang, Weihao Yu, Yujun Shi, Francis EH Tay, Jiashi Feng, and Shuicheng Yan. Tokens-to-token ViT: Training vision transformers from scratch on imagenet. In Proc. ICCV, 2021. 2
  82. 82.Joe Yue-Hei Ng, Matthew Hausknecht, Sudheendra Vijayanarasimhan, Oriol Vinyals, Rajat Monga, and George Toderici. Beyond short snippets: Deep networks for video classification. In Proc. CVPR, 2015. 2
  83. 83.Bolei Zhou, Alex Andonian, Aude Oliva, and Antonio Torralba. Temporal relational reasoning in videos. In ECCV, 2018. 2
  84. 84.Xizhou Zhu, Yujie Wang, Jifeng Dai, Lu Yuan, and Yichen Wei. Flow-guided feature aggregation for video object detection. In Proc. ICCV, 2017. 2
  85. 85.Mohammadreza Zolfaghari, Kamaljeet Singh, and Thomas Brox. ECO: efficient convolutional network for online video understanding. In Proc. ECCV, 2018. 2

Citation

MLA
Wu, C.-Y., et al. “MeMViT: Memory-Augmented Multiscale Vision Transformer for Efficient Long-Term Video Recognition”. arXiv, 2022, http://arxiv.org/abs/2201.08383v2.
APA
Wu, C.-Y., Li, Y., Mangalam, K., Fan, H., Xiong, B., Malik, J., & Feichtenhofer, C. (2022). MeMViT: Memory-Augmented Multiscale Vision Transformer for Efficient Long-Term Video Recognition. arXiv. http://arxiv.org/abs/2201.08383v2
Chicago
Wu, C.-Y., Y. Li, K. Mangalam, et al. 2022. “MeMViT: Memory-Augmented Multiscale Vision Transformer for Efficient Long-Term Video Recognition”. arXiv. http://arxiv.org/abs/2201.08383v2.
Harvard
Wu, C.-Y. et al. (2022) “MeMViT: Memory-Augmented Multiscale Vision Transformer for Efficient Long-Term Video Recognition”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2201.08383v2.
Vancouver
1. Wu C-Y, Li Y, Mangalam K, Fan H, Xiong B, Malik J, Feichtenhofer C (2022) MeMViT: Memory-Augmented Multiscale Vision Transformer for Efficient Long-Term Video Recognition. arXiv

BibTeX

@article{wu2022memvit,
  title = {MeMViT: Memory-Augmented Multiscale Vision Transformer for Efficient Long-Term Video Recognition},
  author = {Wu, Chao-Yuan and Li, Yanghao and Mangalam, Karttikeya and Fan, Haoqi and Xiong, Bo and Malik, Jitendra and Feichtenhofer, Christoph},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2201.08383v2},
  eprint = {2201.08383}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE