SVFormer: Semi-supervised Video Transformer for Action Recognition

Zhen XingQi DaiHan HuJingjing ChenZuxuan WuYu-Gang Jiang

article2023CVPR141 citations

Presents SVFormer, a semi-supervised video transformer framework combining an exponential moving average teacher with specialized tube-level token mixing and temporal warping augmentations to substantially improve action recognition performance under extremely limited supervision.

Listen

Online video content continues to expand rapidly, creating substantial demand for accurate automated video understanding. However, training conventional video recognition models requires vast amounts of manually annotated video, which is labor-intensive and costly. Semi-supervised learning addresses this bottleneck by using a small fraction of labeled videos alongside large volumes of unlabeled footage. While vision transformer models have recently outperformed traditional convolutional neural networks in fully supervised settings, their application to semi-supervised video recognition has remained largely unexplored due to challenges in training transformers with limited supervision and the inadequacy of image-based data augmentation techniques for temporal video.

The article introduces and evaluates SVFormer, a transformer-based semi-supervised action recognition framework designed to train video models effectively using minimal labeled data.

The authors develop a framework incorporating a stable teacher-student model architecture to generate reliable pseudo-labels for unlabeled video clips. To tailor semi-supervised learning to video transformers, the approach introduces two specialized data augmentation techniques: Tube TokenMix, which blends video clips at the transformer token level using temporally consistent spatial masks, and Temporal Warping Augmentation, which randomly stretches and pads frame durations to simulate complex motion variations. The method was evaluated through extensive benchmark experiments on three standard datasets (Kinetics-400, UCF-101, and HMDB-51) under constrained labeling conditions, comparing performance against existing convolutional and transformer baselines.

The evaluation yielded several critical findings. First, SVFormer establishes a new state of the art in semi-supervised video recognition, outperforming prior methods on the Kinetics-400 benchmark with only 1% of data labeled by an absolute 15.0% top-1 accuracy margin for the small model (SVFormer-S) and 31.5% for the base model (SVFormer-B). Second, the framework reaches this higher accuracy significantly faster, requiring only 30 training epochs compared to 180 to 600 epochs for previous convolutional approaches. Third, initializing the video transformer with standard image-pretrained weights resolves the common low-data training failures typically seen in transformer models. Finally, ablation studies confirmed that combining token-level tube masking with temporal warping yields clear performance gains over standard image mixing techniques and isolated spatial augmentations.

These findings demonstrate that organizations can deploy high-performing video recognition systems while reducing data labeling costs and computational training budgets. By relying solely on standard video input without requiring complex auxiliary networks or secondary data streams (such as optical flow), SVFormer streamlines the training pipeline and reduces both operational complexity and carbon footprint. For decision-makers, this shifts the paradigm for video analytics from costly full annotation to lightweight semi-supervised transformer training.

Organizations developing video recognition systems should adopt transformer-based semi-supervised workflows and incorporate token-aligned, temporally aware augmentations rather than legacy image-based mixing methods. Teams seeking to deploy this approach should leverage pretrained image representations as model initializers to stabilize training. Future work should evaluate this framework across longer video sequences, specialized industry-specific domains, and real-time operational environments.

The findings are supported by consistent results across three standard benchmarks and thorough component ablations. Limitations include reliance on short video clips and standard classification settings; performance may vary when transitioning to untrimmed, complex real-world footage or when domain shifts occur between labeled and unlabeled samples.

No sufficiently relevant recommendations were found.

Cover for SVFormer: Semi-supervised Video Transformer for Action Recognition

Abstract

Semi-supervised action recognition is a challenging but critical task due to the high cost of video annotations. Existing approaches mainly use convolutional neural networks, yet current revolutionary vision transformer models have been less explored. In this paper, we investigate the use of transformer models under the SSL setting for action recognition. To this end, we introduce SVFormer, which adopts a steady pseudo-labeling framework (i.e., EMA-Teacher) to cope with unlabeled video samples. While a wide range of data augmentations have been shown effective for semi-supervised image classification, they generally produce limited results for video recognition. We therefore introduce a novel augmentation strategy, Tube TokenMix, tailored for video data where video clips are mixed via a mask with consistent masked tokens over the temporal axis. In addition, we propose a temporal warping augmentation to cover the complex temporal variation in videos, which stretches selected frames to various temporal durations in the clip. Extensive experiments on three datasets Kinetics-400, UCF-101, and HMDB-51 verify the advantage of SVFormer. In particular, SVFormer outperforms the state-of-the-art by 31.5% with fewer training epochs under the 1% labeling rate of Kinetics-400. Our method can hopefully serve as a strong benchmark and encourage future search on semi-supervised action recognition with Transformer networks. Code is released at https://github.com/ChenHsing/SVFormer.

Table of Contents

  • 1. Introduction
  • 2. Related Works
  • 3. Method
  • 3.1. Preliminaries of SSL
  • 3.2. Pipeline
  • 3.3. Tube TokenMix
  • 3.4. Training Paradigm
  • 4. Experiment
  • 4.1. Experiment Settings
  • 4.2. Main Results
  • 4.3. Ablation Studies
  • 5. Conclusion
  • References

Knowls

  1. Knowl 1 — SVFormer Framework for Semi-Supervised Action Recognition

    model/method

    SVFormer is a semi-supervised video action recognition framework built upon video Vision Transformers (such as TimeSformer with divided space-time attention, initialized with ImageNet pre-trained weights). Pre-training on ImageNet provides critical inductive bias that enables transformers to generalize effectively in low-annotation video regimes.

    The framework consists of:

    1. A student transformer network Fs\mathcal{F}_s parameterized by θs\theta_s, and an Exponential Moving Average (EMA) teacher network Ft\mathcal{F}_t parameterized by θt\theta_t. The teacher parameters are updated smoothly via θt←mθt+(1−m)θs\theta_t \leftarrow m \theta_t + (1 - m) \theta_s with momentum coefficient m∈[0,1)m \in [0, 1) to stabilize pseudo-labels and mitigate model collapse.
    2. Supervised training on labeled video clips DL={(xl,yl)}l=1NL\mathcal{D}_L = \{(x_l, y_l)\}_{l=1}^{N_L} using cross-entropy loss.
    3. Unsupervised pseudo-label consistency on unlabeled video clips DU={xu}u=1NU\mathcal{D}_U = \{x_u\}_{u=1}^{N_U}. For each unlabeled clip xux_u, a weakly augmented view xw=Aweak(xu)x_w = \mathcal{A}_{weak}(x_u) (e.g., random horizontal flipping, random scaling, and random cropping) is fed into the teacher model Ft\mathcal{F}_t. When the maximum predicted class probability exceeds a confidence threshold δ\delta, the generated pseudo-label supervises the student model's output on a strongly augmented view xs=Astrong(xu)x_s = \mathcal{A}_{strong}(x_u).
    4. Tube TokenMix (TTMix) consistency regularization, which blends pairs of unlabeled video representations at the token level using temporally continuous tubular masks alongside Temporal Warping Augmentation (TWAug).
  2. Knowl 2 — Tube TokenMix (TTMix) Augmentation

    model/method

    Tube TokenMix (TTMix) is a token-level data augmentation strategy tailored for video Vision Transformers in semi-supervised learning. Unlike pixel-level mixing techniques (such as Mixup or CutMix), TTMix operates on the feature tokens produced after spatial-temporal patch embedding.

    Given two unlabeled video clips xa,xb∈RH×W×T×Cx_a, x_b \in \mathbb{R}^{H \times W \times T \times C} (where HH and WW denote spatial token grid height and width, TT is the temporal clip length, and CC is channel dimension), TTMix blends them using a binary tube mask M∈{0,1}H×W×T\mathbf{M} \in \{0, 1\}^{H \times W \times T}:

    xmix=Astrong(xa)⊙M+Astrong(xb)⊙(1−M)x_{mix} = \mathcal{A}_{strong}(x_a) \odot \mathbf{M} + \mathcal{A}_{strong}(x_b) \odot (\mathbf{1} - \mathbf{M})

    where ⊙\odot denotes element-wise multiplication and 1\mathbf{1} is an all-ones tensor.

    In Tube TokenMix, every temporal frame shares the identical 2D spatial token mask matrix along the temporal axis TT, forming a 3D tube through time. This ensures temporal consistency within each spatial patch tube, preventing information leakage across adjacent frames at identical spatial coordinates and maintaining temporal coherence.

    Alternative mask strategies include:

    • Rand TokenMix: Masked tokens are randomly selected independently across the full H×W×TH \times W \times T volume.
    • Frame TokenMix: Entire temporal frames are randomly selected from the TT frames and fully masked across all spatial locations.

    The pseudo-label y^mix\hat{y}_{mix} for the mixed clip is formed by linearly interpolating the teacher model predictions y^a\hat{y}_a and y^b\hat{y}_b of the two inputs using the mixing ratio λ\lambda (the fraction of ones in M\mathbf{M}, sampled from a Beta distribution Beta(α,α)\text{Beta}(\alpha, \alpha) with α=10\alpha = 10):

    y^mix=λ⋅y^a+(1−λ)⋅y^b\hat{y}_{mix} = \lambda \cdot \hat{y}_a + (1 - \lambda) \cdot \hat{y}_b

  3. Knowl 3 — Temporal Warping Augmentation (TWAug)

    model/method

    Temporal Warping Augmentation (TWAug) is a video augmentation method designed to capture complex, non-uniform temporal dynamics in human actions by distorting the temporal duration of individual frames in a video clip.

    Given an extracted video clip of TT frames (e.g., T=8T = 8), TWAug operates as follows:

    1. Randomly decide whether to keep all TT frames or subsample a small subset of K<TK < T frames (e.g., K∈{2,4}K \in \{2, 4\}).
    2. Place the selected KK frames at their respective temporal indices and mask out all unselected frame positions.
    3. Pad the masked frame positions with randomly selected neighboring visible (unmasked) frames while strictly preserving chronological frame ordering.

    For example, starting from an 8-frame clip [1,2,3,4,5,6,7,8][1, 2, 3, 4, 5, 6, 7, 8]:

    • Subsampling 2 frames {3,6}\{3, 6\} yields a warped sequence such as [3,3,3,3,6,6,6,6][3, 3, 3, 3, 6, 6, 6, 6].
    • Subsampling 4 frames {1,5,7,8}\{1, 5, 7, 8\} yields a warped sequence such as [1,1,1,5,5,7,7,8][1, 1, 1, 5, 5, 7, 7, 8].

    TWAug acts as a strong temporal augmentation. Within the Tube TokenMix pipeline, one input clip undergoes standard spatial augmentation (e.g., AutoAugment or Dropout) while the paired clip undergoes TWAug prior to token mixing.

  4. Knowl 4 — Consistency Loss Algorithm for Tube TokenMix

    algorithm

    The Tube TokenMix consistency regularization algorithm processes batches of unlabeled video clips by shuffling, applying spatial and temporal warping augmentations, generating interpolated teacher pseudo-labels with threshold filtering, and calculating the mean squared error loss on student predictions.

    Input: Unlabeled clip batch xax_a, Teacher model Ft\mathcal{F}_t, Student model Fs\mathcal{F}_s, Tube TokenMask M\mathbf{M}, Mask ratio λ∈[0,1]\lambda \in [0, 1], Confidence threshold δ∈[0,1]\delta \in [0, 1]
    Output: Consistency loss Lmix\mathcal{L}_{mix}
    xb←shuffle(xa)x_b \leftarrow \text{shuffle}(x_a)
    x^a←spatial_aug(xa)\hat{x}_a \leftarrow \text{spatial\_aug}(x_a)
    x^b←temporal_aug(xb)\hat{x}_b \leftarrow \text{temporal\_aug}(x_b)
    y^a←stop_gradient(Ft(xa))\hat{y}_a \leftarrow \text{stop\_gradient}(\mathcal{F}_t(x_a))
    y^b←stop_gradient(Ft(xb))\hat{y}_b \leftarrow \text{stop\_gradient}(\mathcal{F}_t(x_b))
    ca←max⁡iy^a[i]c_a \leftarrow \max_i \hat{y}_a[i]
    cb←max⁡iy^b[i]c_b \leftarrow \max_i \hat{y}_b[i]
    xmix←x^a⊙M+x^b⊙(1−M)x_{mix} \leftarrow \hat{x}_a \odot \mathbf{M} + \hat{x}_b \odot (\mathbf{1} - \mathbf{M})
    y^mix←y^a⋅λ+y^b⋅(1−λ)\hat{y}_{mix} \leftarrow \hat{y}_a \cdot \lambda + \hat{y}_b \cdot (1 - \lambda)
    cmix←ca⋅λ+cb⋅(1−λ)c_{mix} \leftarrow c_a \cdot \lambda + c_b \cdot (1 - \lambda)
    q←mean(cmix≥δ)q \leftarrow \text{mean}(c_{mix} \ge \delta)
    ymix←Fs(xmix)y_{mix} \leftarrow \mathcal{F}_s(x_{mix})
    Lmix←q⋅∥ymix−y^mix∥22\mathcal{L}_{mix} \leftarrow q \cdot \|y_{mix} - \hat{y}_{mix}\|_2^2
    return Lmix\mathcal{L}_{mix}

    Here, spatial_aug\text{spatial\_aug} applies spatial transformations (AutoAugment, Dropout), temporal_aug\text{temporal\_aug} applies Temporal Warping Augmentation (TWAug), and stop_gradient\text{stop\_gradient} detaches teacher outputs from backpropagation.

  5. Knowl 5 — SVFormer Loss Formulation

    equation

    The overall training objective of SVFormer integrates three loss components: supervised classification loss Ls\mathcal{L}_s, unsupervised pseudo-label consistency loss Lun\mathcal{L}_{un}, and Tube TokenMix consistency loss Lmix\mathcal{L}_{mix}:

    Lall=Ls+γ1Lun+γ2Lmix\mathcal{L}_{all} = \mathcal{L}_s + \gamma_1 \mathcal{L}_{un} + \gamma_2 \mathcal{L}_{mix}

    where γ1\gamma_1 and γ2\gamma_2 are balancing hyperparameters (default γ1=2,γ2=2\gamma_1 = 2, \gamma_2 = 2).

    1. Supervised Loss on labeled set DL={(xl,yl)}l=1NL\mathcal{D}_L = \{(x_l, y_l)\}_{l=1}^{N_L}: Ls=1NL∑l=1NLH(Fs(xl),yl)\mathcal{L}_s = \frac{1}{N_L} \sum_{l=1}^{N_L} \mathcal{H}(\mathcal{F}_s(x_l), y_l) where Fs(⋅)\mathcal{F}_s(\cdot) is the student model prediction and H\mathcal{H} is standard cross-entropy.

    2. Unsupervised Pseudo-Label Loss on unlabeled set DU={xu}u=1NU\mathcal{D}_U = \{x_u\}_{u=1}^{N_U}: Lun=1NU∑u=1NUI(max⁡(Ft(xw))>δ)H(Fs(xs),y^w)\mathcal{L}_{un} = \frac{1}{N_U} \sum_{u=1}^{N_U} \mathbb{I}(\max(\mathcal{F}_t(x_w)) > \delta) \mathcal{H}(\mathcal{F}_s(x_s), \hat{y}_w) where xw=Aweak(xu)x_w = \mathcal{A}_{weak}(x_u), xs=Astrong(xu)x_s = \mathcal{A}_{strong}(x_u), Ft\mathcal{F}_t is the EMA teacher model, y^w=arg⁡max⁡(Ft(xw))\hat{y}_w = \arg\max(\mathcal{F}_t(x_w)), δ\delta is the confidence threshold, and I(⋅)\mathbb{I}(\cdot) is the indicator function.

    3. Tube TokenMix Consistency Loss on mixed samples: Lmix=1Nm∑k=1Nm(y^mix,k−ymix,k)2\mathcal{L}_{mix} = \frac{1}{N_m} \sum_{k=1}^{N_m} (\hat{y}_{mix, k} - y_{mix, k})^2 where NmN_m is the number of mixed samples, ymix=Fs(xmix)y_{mix} = \mathcal{F}_s(x_{mix}) is the student prediction on the token-mixed clip, and y^mix=λy^a+(1−λ)y^b\hat{y}_{mix} = \lambda \hat{y}_a + (1 - \lambda) \hat{y}_b is the interpolated teacher pseudo-label.

  6. Knowl 6 — Semi-Supervised Action Recognition Benchmark on UCF-101 and Kinetics-400

    data/table

    Semi-supervised action recognition performance comparison on UCF-101 and Kinetics-400 under 1% and 10% labeling ratios. Top-1 accuracy (%) is reported.

    Method Backbone Input Epoch UCF-101 (1%) UCF-101 (10%) Kinetics-400 (1%) Kinetics-400 (10%)
    Supervised 3D-ResNet-50 V 200 6.5 32.4 4.4 36.2
    Supervised ViT-S V 30 12.7 62.5 19.9 56.6
    FixMatch SlowFast-R50 V 200 16.1 55.1 10.1 49.4
    VideoSSL 3D-ResNet-18 V - - 42.0 - 33.8
    TCL TSM-ResNet-18 V 400 - - 8.5 -
    ActorCutMix R(2+1)D-34 V 600 - 53.0 9.02 -
    MvPL 3D-ResNet-50 V+F+G 600 22.8 80.5 17.0 58.2
    CMPL R50 + R50-1/4 V 200 25.1 79.1 17.6 58.4
    LTG 3D-ResNet-18 V+G 180/360 - 62.4 9.8 43.8
    TACL 3D-ResNet-50 V 200 - 55.6 - -
    L2A 3D-ResNet-18 V 400 - 60.1 - -
    SVFormer-S (Ours) ViT-S V 30 31.4 79.1 32.6 61.6
    SVFormer-B (Ours) ViT-B V 30 46.3 86.7 49.1 69.4

    Input modalities: "V" denotes raw RGB video, "F" optical flow, and "G" temporal gradient. SVFormer-S trained for 30 epochs outperforms previous state-of-the-art methods (e.g., CMPL) by 6.3% on UCF-101 (1%) and by 15.0% on Kinetics-400 (1%). SVFormer-B scales to 46.3% on UCF-101 (1%) and 49.1% on Kinetics-400 (1%), and reaches 69.4% on Kinetics-400 (10%), approaching fully supervised TimeSformer accuracy (77.9%).

  7. Knowl 7 — Semi-Supervised Action Recognition Benchmark on HMDB-51

    data/table

    Semi-supervised action recognition performance comparison on HMDB-51 across 40%, 50%, and 60% labeling ratios. Top-1 accuracy (%) is reported.

    Method Backbone Input HMDB-51 (40%) HMDB-51 (50%) HMDB-51 (60%)
    VideoSSL 3D-R18 V 32.7 36.2 37.0
    ActorCutMix R(2+1)D-34 V 32.9 38.2 38.9
    MvPL 3D-R18 V+F+G 30.5 33.9 35.8
    LTG 3D-R18 V+G 46.5 48.4 49.7
    TACL 3D-R18 V 38.7 40.2 41.7
    L2A 3D-R18 V 42.1 46.3 47.1
    SVFormer-S (Ours) ViT-S V 56.2 58.2 59.7
    SVFormer-B (Ours) ViT-B V 61.6 64.4 68.2

    Input modalities: "V" denotes raw RGB video, "F" optical flow, and "G" temporal gradient. "3D-R18" denotes 3D-ResNet-18. SVFormer-S improves over the best prior multi-modal baseline (LTG with RGB + temporal gradient) by ~10% Top-1 accuracy across all splits using solely RGB video. SVFormer-B achieves further gains of 5.4% to 8.5% over SVFormer-S.

  8. Knowl 8 — Ablation of Mixing Strategies and Token Masking in Video Transformers

    data/table

    Evaluation of various pixel-level and token-level mixing augmentation strategies using SVFormer-S on UCF-101 and Kinetics-400 with 1% labeled data. Top-1 and Top-5 accuracy (%) are reported.

    Method UCF-101 (1%) Kinetics-400 (1%)
    Top-1 Top-5 Top-1 Top-5
    Baseline 26.1 48.9 23.6 47.7
    CutMix 28.7 51.3 28.6 53.7
    Mixup 29.8 53.0 29.3 55.1
    PixMix 29.7 52.4 29.6 55.8
    Frame TokenMix 29.8 54.2 26.3 50.0
    Rand TokenMix 30.3 55.3 28.8 54.2
    Tube TokenMix 31.4 56.9 32.6 59.0

    Token-level mixing methods operate naturally on Vision Transformer patch tokens and outperform pixel-level methods (CutMix, Mixup, PixMix). Frame TokenMix yields inferior results on Kinetics-400 (26.3% Top-1) because masking entire frames disrupts temporal sequence modeling. Tube TokenMix achieves the highest performance (31.4% on UCF-101 and 32.6% on Kinetics-400) by preserving temporal coherence along each spatial tube.

  9. Knowl 9 — Ablation of Spatial and Temporal Warping Augmentations

    data/table

    Ablation study of strong spatial augmentation (AutoAugment and Dropout) and Temporal Warping Augmentation (TWAug) within SVFormer-S on UCF-101 (1%) and Kinetics-400 (1%). Top-1 and Top-5 accuracy (%) are reported.

    Spatial Temporal UCF-101 (1%) Kinetics-400 (1%)
    Top-1 Top-5 Top-1 Top-5
    29.5 52.1 28.5 54.2
    ✓ 29.9 53.2 30.3 55.9
    ✓ 30.2 55.8 30.8 56.4
    ✓ ✓ 31.4 56.9 32.6 59.0

    Both spatial augmentation alone and temporal warping augmentation alone improve performance over the unaugmented baseline (Top-1 on Kinetics-400 rises from 28.5% to 30.3% and 30.8%, respectively). Combining both spatial and temporal warping augmentations achieves the best performance (31.4% on UCF-101 and 32.6% on Kinetics-400), showing that spatial and temporal augmentations provide complementary regularization.

  10. Knowl 10 — Hyperparameter and Framework Analysis for SVFormer

    empirical result

    Empirical sensitivity analysis of SVFormer-S on Kinetics-400 (1% labeled) and UCF-101 (1% labeled) establishes the following optimal hyperparameter behaviors:

    1. SSL Framework: The EMA-Teacher framework substantially outperforms FixMatch without EMA. On UCF-101 (1%), EMA-Teacher achieves 31.4% Top-1 (vs. 25.1% for FixMatch and 12.7% for baseline). On Kinetics-400 (1%), EMA-Teacher achieves 32.6% Top-1 (vs. 28.2% for FixMatch and 19.9% for baseline). EMA provides parameter stability that prevents model collapse when labeled samples are extremely sparse.
    2. Confidence Threshold δ\delta: Swept across δ∈{0.1,0.3,0.5,0.7,0.9}\delta \in \{0.1, 0.3, 0.5, 0.7, 0.9\} on Kinetics-400 (1%), the optimal threshold is δ=0.3\delta = 0.3 (32.6% Top-1), whereas higher thresholds filter out too many valid training signals (dropping to 29.5% at δ=0.9\delta = 0.9).
    3. Unlabeled-to-Labeled Batch Ratio Bu/BlB_u / B_l: Swept across {1,2,3,5,7}\{1, 2, 3, 5, 7\} with labeled batch size Bl=1B_l = 1, Top-1 accuracy peaks at Bu=5B_u = 5 (32.6% Top-1, compared to 27.2% at Bu=1B_u = 1 and 31.5% at Bu=7B_u = 7).
    4. EMA Momentum mm: Tested across m∈{0.9,0.99,0.999,0.9999}m \in \{0.9, 0.99, 0.999, 0.9999\}, m=0.99m = 0.99 yields the highest accuracy (32.6% vs. 30.9% at 0.90.9 and 30.3% at 0.99990.9999).
    5. Loss Weights γ1,γ2\gamma_1, \gamma_2: Setting γ1=γ2=2\gamma_1 = \gamma_2 = 2 provides the best performance (32.6% Top-1 vs. 31.4% at weight 1 and 30.9% at weight 5).
    6. Inference Scheme: Clip-based sampling (8 frames×rate 88 \text{ frames} \times \text{rate } 8 with 5×35 \times 3 crops) yields 31.4% (UCF-101 1%) and 32.6% (Kinetics-400 1%), outperforming video-based sparse sampling (8×328 \times 32 with 1×31 \times 3 crops, yielding 29.3% and 31.0%).

Coverage note — None was omitted; all contributed models, augmentations, algorithms, loss formulations, benchmark comparisons, and ablation studies are fully covered.

References

  1. 1.Inigo Alonso, Alberto Sabater, David Ferstl, Luis Montesano, and Ana C Murillo. Semi-supervised semantic segmentation with pixel-level contrastive learning from a class-wise memory bank. In ICCV, 2021.
  2. 2.Anurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun, Mario Lućić, and Cordelia Schmid. Vivit: A video vision transformer. In ICCV, 2021.
  3. 3.Steven S. Beauchemin and John L. Barron. The computation of optical flow. ACM computing surveys (CSUR), 1995.
  4. 4.Gedas Bertasius, Heng Wang, and Lorenzo Torresani. Is space-time attention all you need for video understanding? In ICML, 2021.
  5. 5.David Berthelot, Nicholas Carlini, Ian Goodfellow, Nicolas Papernot, Avital Oliver, and Colin A Raffel. Mixmatch: A holistic approach to semi-supervised learning. In NeurIPS, 2019.
  6. 6.Zhaowei Cai, Avinash Ravichandran, Paolo Favaro, Manchen Wang, Davide Modolo, Rahul Bhotika, Zhuowen Tu, and Stefano Soatto. Semi-supervised vision transformers at scale. In NeurIPS, 2022.
  7. 7.Joao Carreira and Andrew Zisserman. Quo vadis, action recognition? a new model and the kinetics dataset. In CVPR, 2017.
  8. 8.Xiaokang Chen, Yuhui Yuan, Gang Zeng, and Jingdong Wang. Semi-supervised semantic segmentation with cross pseudo supervision. In CVPR, 2021.
  9. 9.Hsin-Ping Chou, Shih-Chieh Chang, Jia-Yu Pan, Wei Wei, and Da-Cheng Juan. Remix: rebalanced mixup. In ECCV, 2020.
  10. 10.Ekin D Cubuk, Barret Zoph, Dandelion Mane, Vijay Vasudevan, and Quoc V Le. Autoaugment: Learning augmentation strategies from data. In CVPR, 2019.
  11. 11.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In CVPR, 2009.
  12. 12.Terrance DeVries and Graham W Taylor. Improved regularization of convolutional neural networks with cutout. arXiv preprint arXiv:1708.04552, 2017.
  13. 13.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. In ICLR, 2021.
  14. 14.Haoqi Fan, Bo Xiong, Karttikeya Mangalam, Yanghao Li, Zhicheng Yan, Jitendra Malik, and Christoph Feichtenhofer. Multiscale vision transformers. In ICCV, 2021.
  15. 15.Christoph Feichtenhofer. X3d: Expanding architectures for efficient video recognition. In CVPR, 2020.
  16. 16.Christoph Feichtenhofer, Haoqi Fan, Yanghao Li, and Kaiming He. Masked autoencoders as spatiotemporal learners. In NeurIPS, 2022.
  17. 17.Christoph Feichtenhofer, Haoqi Fan, Jitendra Malik, and Kaiming He. Slowfast networks for video recognition. In ICCV, 2019.
  18. 18.Christoph Feichtenhofer, Haoqi Fan, Bo Xiong, Ross Girshick, and Kaiming He. A large-scale study on unsupervised spatiotemporal representation learning. In CVPR, 2021.
  19. 19.Geoff French, Avital Oliver, and Tim Salimans. Milking cowmask for semi-supervised image classification. arXiv preprint arXiv:2003.12022, 2020.
  20. 20.Shreyank N Gowda, Marcus Rohrbach, Frank Keller, and Laura Sevilla-Lara. Learn2augment: Learning to composite videos for data augmentation in action recognition. In ECCV, 2022.
  21. 21.Jean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec, Pierre Richemond, Elena Buchatskaya, Carl Doersch, Bernardo Avila Pires, Zhaohan Guo, Mohammad Gheshlaghi Azar, et al. Bootstrap your own latent-a new approach to self-supervised learning. NeurIPS, 2020.
  22. 22.Kensho Hara, Hirokatsu Kataoka, and Yutaka Satoh. Learning spatio-temporal features with 3d residual networks for action recognition. In ICCVW, 2017.
  23. 23.Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick. Momentum contrast for unsupervised visual representation learning. In CVPR, 2020.
  24. 24.Dan Hendrycks, Andy Zou, Mantas Mazeika, Leonard Tang, Bo Li, Dawn Song, and Jacob Steinhardt. Pixmix: Dreamlike pictures comprehensively improve safety measures. In CVPR, 2022.
  25. 25.Longlong Jing, Toufiq Parag, Zhe Wu, Yingli Tian, and Hongcheng Wang. Videossl: Semi-supervised learning for video classification. In WACV, 2021.
  26. 26.Hildegard Kuehne, Hueihan Jhuang, Estíbaliz Garrote, Tomaso Poggio, and Thomas Serre. Hmdb: a large video database for human motion recognition. In ICCV, 2011.
  27. 27.Hengduo Li, Zuxuan Wu, Abhinav Shrivastava, and Larry S Davis. Rethinking pseudo labels for semi-supervised object detection. In AAAI, 2022.
  28. 28.Kunchang Li, Yali Wang, Peng Gao, Guanglu Song, Yu Liu, Hongsheng Li, and Yu Qiao. Uniformer: Unified transformer for efficient spatiotemporal representation learning. In ICLR, 2022.
  29. 29.Wenxu Li, Gang Pan, Chen Wang, Zhen Xing, and Zhenjun Han. From coarse to fine: Hierarchical structure-aware video summarization. ACM TOMM), 2022.
  30. 30.Zhixin Ling, Zhen Xing, Xiangdong Zhou, Manliang Cao, and Guichun Zhou. Panoswin: a pano-style swin transformer for panorama understanding. In CVPR, 2023.
  31. 31.Jihao Liu, Boxiao Liu, Hang Zhou, Hongsheng Li, and Yu Liu. Tokenmix: Rethinking image mixing for data augmentation in vision transformers. In ECCV, 2022.
  32. 32.Yen-Cheng Liu, Chih-Yao Ma, Zijian He, Chia-Wen Kuo, Kan Chen, Peizhao Zhang, Bichen Wu, Zsolt Kira, and Peter Vajda. Unbiased teacher for semi-supervised object detection. In ICLR, 2021.
  33. 33.Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. In ICCV, 2021.
  34. 34.Ze Liu, Jia Ning, Yue Cao, Yixuan Wei, Zheng Zhang, Stephen Lin, and Han Hu. Video swin transformer. In CVPR, 2022.
  35. 35.Puneet Mangla, Nupur Kumari, Abhishek Sinha, Mayank Singh, Balaji Krishnamurthy, and Vineeth N Balasubramanian. Charting the right manifold: Manifold mixup for few-shot learning. In WACV, 2020.
  36. 36.Daniel Neimark, Omri Bar, Maya Zohar, and Dotan Asselmann. Video transformer network. In ICCVW, 2021.
  37. 37.Farrukh Rahman, Omer Mubarek, and Zsolt Kira. On the surprising effectiveness of transformers in low-labeled video recognition. In NeurIPSW, 2022.
  38. 38.Baifeng Shi, Qi Dai, Judy Hoffman, Kate Saenko, Trevor Darrell, and Huijuan Xu. Temporal action detection with multi-level supervision. In ICCV, 2021.
  39. 39.Baifeng Shi, Qi Dai, Yadong Mu, and Jingdong Wang. Weakly-supervised action localization by generative attention modeling. In CVPR, 2020.
  40. 40.Ankit Singh, Omprakash Chakraborty, Ashutosh Varshney, Rameswar Panda, Rogerio Feris, Kate Saenko, and Abir Das. Semi-supervised action recognition with temporal contrastive learning. In CVPR, 2021.
  41. 41.Kihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang, Han Zhang, Colin A Raffel, Ekin Dogus Cubuk, Alexey Kurakin, and Chun-Liang Li. Fixmatch: Simplifying semi-supervised learning with consistency and confidence. In NeurIPS, 2020.
  42. 42.Khurram Soomro, Amir Roshan Zamir, and Mubarak Shah. Ucf101: A dataset of 101 human actions classes from videos in the wild. arXiv preprint arXiv:1212.0402, 2012.
  43. 43.Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. Dropout: a simple way to prevent neural networks from overfitting. The journal of machine learning research, 2014.
  44. 44.Antti Tarvainen and Harri Valpola. Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results. In NeurIPS, volume 30, 2017.
  45. 45.Rui Tian, Zuxuan Wu, Qi Dai, Han Hu, Yu Qiao, and Yu-Gang Jiang. Resformer: Scaling vits with multi-resolution training. In CVPR, 2023.
  46. 46.Anyang Tong, Chao Tang, and Wenjian Wang. Semi-supervised action recognition from temporal augmentation using curriculum learning. IEEE TCSVT, 2022.
  47. 47.Zhan Tong, Yibing Song, Jue Wang, and Limin Wang. Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training. In NeurIPS, 2022.
  48. 48.Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Hervé Jégou. Training data-efficient image transformers & distillation through attention. In ICML, 2021.
  49. 49.Vikas Verma, Kenji Kawaguchi, Alex Lamb, Juho Kannala, Yoshua Bengio, and David Lopez-Paz. Interpolation consistency training for semi-supervised learning. In IJCAI, 2019.
  50. 50.Limin Wang, Yuanjun Xiong, Zhe Wang, Yu Qiao, Dahua Lin, Xiaoou Tang, and Luc Van Gool. Temporal segment networks: Towards good practices for deep action recognition. In ECCV, 2016.
  51. 51.Rui Wang, Dongdong Chen, Zuxuan Wu, Yinpeng Chen, Xiyang Dai, Mengchen Liu, Yu-Gang Jiang, Luowei Zhou, and Lu Yuan. Bevt: Bert pretraining of video transformers. In CVPR, 2022.
  52. 52.Rui Wang, Dongdong Chen, Zuxuan Wu, Yinpeng Chen, Xiyang Dai, Mengchen Liu, Lu Yuan, and Yu-Gang Jiang. Masked video distillation: Rethinking masked feature modeling for self-supervised video representation learning. In CVPR, 2023.
  53. 53.Rui Wang, Zuxuan Wu, Dongdong Chen, Yinpeng Chen, Xiyang Dai, Mengchen Liu, Luowei Zhou, Lu Yuan, and Yu-Gang Jiang. Video mobile-former: Video recognition with efficient global spatial-temporal modeling. arXiv preprint arXiv:2208.12257, 2022.
  54. 54.Zejia Weng, Xitong Yang, Ang Li, Zuxuan Wu, and Yu-Gang Jiang. Semi-supervised vision transformers. In ECCV, 2022.
  55. 55.Yuan Wu, Diana Inkpen, and Ahmed El-Roby. Dual mixup regularized learning for adversarial domain adaptation. In ECCV, 2020.
  56. 56.Zuxuan Wu, Hengduo Li, Yingbin Zheng, Caiming Xiong, Yu-Gang Jiang, and Davis Larry S. A coarse-to-fine framework for resource efficient video recognition. IJCV, 2021.
  57. 57.Junfei Xiao, Longlong Jing, Lin Zhang, Ju He, Qi She, Zongwei Zhou, Alan Yuille, and Yingwei Li. Learning from temporal gradient for semi-supervised action recognition. In CVPR, 2022.
  58. 58.Zhenda Xie, Zheng Zhang, Yue Cao, Yutong Lin, Jianmin Bao, Zhuliang Yao, Qi Dai, and Han Hu. Simmim: A simple framework for masked image modeling. In CVPR, 2022.
  59. 59.Zhenda Xie, Zheng Zhang, Yue Cao, Yutong Lin, Yixuan Wei, Qi Dai, and Han Hu. On data scaling in masked image modeling. In CVPR, 2023.
  60. 60.Zhen Xing, Yijiang Chen, Zhixin Ling, Xiangdong Zhou, and Yu Xiang. Few-shot single-view 3d reconstruction with memory prior contrastive network. In ECCV, 2022.
  61. 61.Zhen Xing, Hengduo Li, Zuxuan Wu, and Yu-Gang Jiang. Semi-supervised single-view 3d reconstruction via prototype shape priors. In ECCV, 2022.
  62. 62.Bo Xiong, Haoqi Fan, Kristen Grauman, and Christoph Feichtenhofer. Multiview pseudo-labeling for semi-supervised learning from video. In ICCV, 2021.
  63. 63.Minghao Xu, Jian Zhang, Bingbing Ni, Teng Li, Chengjie Wang, Qi Tian, and Wenjun Zhang. Adversarial domain adaptation with domain mixup. In AAAI, 2020.
  64. 64.Yinghao Xu, Fangyun Wei, Xiao Sun, Ceyuan Yang, Yujun Shen, Bo Dai, Bolei Zhou, and Stephen Lin. Cross-model pseudo-labeling for semi-supervised action recognition. In CVPR, 2022.
  65. 65.Ceyuan Yang, Yinghao Xu, Bo Dai, and Bolei Zhou. Video representation learning with visual tempo consistency. arXiv preprint arXiv:2006.15489, 2020.
  66. 66.Sangdoo Yun, Dongyoon Han, Seong Joon Oh, Sanghyuk Chun, Junsuk Choe, and Youngjoon Yoo. Cutmix: Regularization strategy to train strong classifiers with localizable features. In ICCV, 2019.
  67. 67.Hongyi Zhang, Moustapha Cisse, Yann N Dauphin, and David Lopez-Paz. mixup: Beyond empirical risk minimization. In ICLR, 2018.
  68. 68.Xiaosong Zhang, Yunjie Tian, Lingxi Xie, Wei Huang, Qi Dai, Qixiang Ye, and Qi Tian. Hivit: A simpler and more efficient design of hierarchical vision transformer. In ICLR, 2023.
  69. 69.Xing Zhang, Zuxuan Wu, Zejia Weng, Huazhu Fu, Jingjing Chen, Yu-Gang Jiang, and Larry S Davis. Videolt: Large-scale long-tailed video recognition. In ICCV, 2021.
  70. 70.Linchao Zhu and Yi Yang. Actbert: Learning global-local video-text representations. In CVPR, 2020.
  71. 71.Yuliang Zou, Jinwoo Choi, Qitong Wang, and Jia-Bin Huang. Learning representational invariances for data-efficient action recognition. arXiv preprint arXiv:2103.16565, 2021.

Citation

MLA
Xing, Z., et al. “SVFormer: Semi-supervised Video Transformer for Action Recognition”. arXiv, 2022, http://arxiv.org/abs/2211.13222v2.
APA
Xing, Z., Dai, Q., Hu, H., Chen, J., Wu, Z., & Jiang, Y.-G. (2022). SVFormer: Semi-supervised Video Transformer for Action Recognition. arXiv. http://arxiv.org/abs/2211.13222v2
Chicago
Xing, Z., Q. Dai, H. Hu, J. Chen, Z. Wu, and Y.-G. Jiang. 2022. “SVFormer: Semi-supervised Video Transformer for Action Recognition”. arXiv. http://arxiv.org/abs/2211.13222v2.
Harvard
Xing, Z. et al. (2022) “SVFormer: Semi-supervised Video Transformer for Action Recognition”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2211.13222v2.
Vancouver
1. Xing Z, Dai Q, Hu H, Chen J, Wu Z, Jiang Y-G (2022) SVFormer: Semi-supervised Video Transformer for Action Recognition. arXiv

BibTeX

@article{xing2022svformer,
  title = {SVFormer: Semi-supervised Video Transformer for Action Recognition},
  author = {Xing, Zhen and Dai, Qi and Hu, Han and Chen, Jingjing and Wu, Zuxuan and Jiang, Yu-Gang},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2211.13222v2},
  eprint = {2211.13222}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE