A Closer Look at Spatiotemporal Convolutions for Action Recognition

Du TranHeng WangLorenzo TorresaniJamie RayYann LeCunManohar Paluri

article2017CVPR3,724 citations

Introduces the R(2+1)D architecture, which factorizes 3D spatiotemporal convolutions into separate 2D spatial and 1D temporal operations to surpass conventional 3D CNNs across major action recognition benchmarks.

Listen

Recent advances in still-image recognition have relied on deep residual networks, yet video action recognition has seen more modest gains from similar techniques despite the importance of motion. Two-dimensional convolutional networks applied frame-by-frame remain competitive on large benchmarks, raising the question of whether explicit temporal modeling is necessary. This work examines that question by comparing multiple forms of spatiotemporal convolution inside residual architectures on major action datasets.

The study evaluates five convolutional variantspure 2D, pure 3D, mixed 3D-2D, reversed mixed, and a factorized (2+1)D blockon Kinetics and Sports-1M, with transfer tests on UCF101 and HMDB51. All models were trained from scratch or fine-tuned under controlled conditions using 18- and 34-layer ResNets, with consistent input sizes, optimization schedules, and evaluation protocols that measure both clip-level and video-level accuracy.

The clearest result is that R(2+1)D, which replaces every 3D filter with a spatial 2D convolution followed by a temporal 1D convolution, delivers the highest accuracy across all four benchmarks while using the same parameter count as full 3D networks. On Kinetics it improves clip accuracy by 35 points over both 3D and mixed-convolution baselines; on Sports-1M it exceeds the previous best published result by roughly 9 points in clip accuracy. The factorization also produces lower training and test error than equivalent 3D networks, with the gap widening at greater depth, indicating easier optimization and the benefit of an extra nonlinearity between spatial and temporal stages. Mixed-convolution networks that apply 3D layers only in early stages match full 3D performance at one-third the parameter cost, confirming that motion modeling matters most at lower levels.

These outcomes show that temporal reasoning remains valuable for action recognition once the right decomposition is used, and that the performance gap between 2D and 3D models can be closed or reversed without increasing model size. The findings support wider adoption of factorized spatiotemporal blocks in video pipelines where both accuracy and training stability matter.

Further gains will likely require systematic architecture search beyond the homogeneous ResNet backbone tested here, together with larger-scale pretraining and more efficient optical-flow methods. The study is limited to residual networks and four public datasets; results may shift under different backbone families or domain-specific video distributions.

arXiv: 1711.11248
Cover for A Closer Look at Spatiotemporal Convolutions for Action Recognition

Abstract

In this paper we discuss several forms of spatiotemporal convolutions for video analysis and study their effects on action recognition. Our motivation stems from the observation that 2D CNNs applied to individual frames of the video have remained solid performers in action recognition. In this work we empirically demonstrate the accuracy advantages of 3D CNNs over 2D CNNs within the framework of residual learning. Furthermore, we show that factorizing the 3D convolutional filters into separate spatial and temporal components yields significantly advantages in accuracy. Our empirical study leads to the design of a new spatiotemporal convolutional block "R(2+1)D" which gives rise to CNNs that achieve results comparable or superior to the state-of-the-art on Sports-1M, Kinetics, UCF101 and HMDB51.

Table of Contents

  • A Closer Look at Spatiotemporal Convolutions for Action Recognition
  • 1. Introduction
  • 2. Related Work
  • 3. Convolutional residual blocks for video
  • 3.1. R2D: 2D convolutions over the entire clip
  • 3.2. f-R2D: 2D convolutions over frames
  • 3.3. R3D: 3D convolutions
  • 3.4. MCx and rMCx: mixed 3D-2D convolutions
  • 3.5. R(2+1)D: (2+1)D convolutions
  • 4. Experiments
  • 4.1. Experimental setup
  • 4.2. Comparison of spatiotemporal convolutions
  • 4.3. Revisiting practices for video-level prediction
  • 4.4. Action recognition with a 34-layer R(2+1)D net
  • 5. Conclusions

Knowls

  1. Knowl 1 — (2+1)D Spatiotemporal Convolution Block

    model/method

    The (2+1)D(2+1)\text{D} convolutional block decomposes a full 3D spatiotemporal convolution into two successive operations: a 2D spatial convolution followed by a 1D temporal convolution.

    Let an input feature tensor to the ii-th residual convolutional block have Ni1N_{i-1} feature channels. Instead of applying NiN_i full 3D convolutional filters of size Ni1×t×d×dN_{i-1} \times t \times d \times d (where tt is the temporal filter extent and d×dd \times d is the spatial filter size), the (2+1)D(2+1)\text{D} block replaces each 3D convolution with:

    1. A 2D spatial convolution consisting of MiM_i filters of size Ni1×1×d×dN_{i-1} \times 1 \times d \times d, projecting the features into an intermediate subspace of dimension MiM_i.
    2. A Batch Normalization layer and a non-linear ReLU activation function.
    3. A 1D temporal convolution consisting of NiN_i filters of size Mi×t×1×1M_i \times t \times 1 \times 1.
    4. A Batch Normalization layer.

    When spatiotemporal downsampling is required (via striding), spatial striding is executed by the 2D convolution and temporal striding is executed by the 1D convolution.

  2. Knowl 2 — Intermediate Channel Dimension for (2+1)D Parameter Matching

    equation

    To ensure that a (2+1)D(2+1)\text{D} convolutional block has approximately the same number of learnable parameters as a standard 3D convolutional block with NiN_i filters of size Ni1×t×d×dN_{i-1} \times t \times d \times d, the intermediate channel dimension MiM_i between the 2D spatial convolution and the 1D temporal convolution is set to:

    Mi=td2Ni1Nid2Ni1+tNiM_i = \left\lfloor \frac{t d^2 N_{i-1} N_i}{d^2 N_{i-1} + t N_i} \right\rfloor

    where:

    • Ni1NN_{i-1} \in \mathbb{N} is the number of input channels to the ii-th block,
    • NiNN_i \in \mathbb{N} is the number of output channels of the ii-th block,
    • dNd \in \mathbb{N} is the spatial height and width of the convolutional filter (e.g., d=3d=3),
    • tNt \in \mathbb{N} is the temporal extent of the filter (e.g., t=3t=3),
    • \lfloor \cdot \rfloor denotes the floor function.

    This formula equates the parameter count of full 3D convolution, Ni1Nitd2N_{i-1} \cdot N_i \cdot t \cdot d^2, with that of the factorized sequence, Ni1Mi1d2+MiNit12=Mi(d2Ni1+tNi)N_{i-1} \cdot M_i \cdot 1 \cdot d^2 + M_i \cdot N_i \cdot t \cdot 1^2 = M_i (d^2 N_{i-1} + t N_i).

  3. Knowl 3 — Family of Spatiotemporal Convolutional Residual Networks

    model/method

    Several structural variants of spatiotemporal residual networks (ResNets) can be constructed to explore trade-offs between spatial and temporal modeling across network depth for an input video clip xR3×L×H×Wx \in \mathbb{R}^{3 \times L \times H \times W} (LL frames of size H×WH \times W with 3 RGB channels):

    • R2D (Clip-based 2D ResNet): Treats the LL frames as input channels (reshaping input to 3L×H×W3L \times H \times W). Filters have size Ni1×d×dN_{i-1} \times d \times d and convolve solely over space. Temporal information is permanently collapsed into 2D spatial channels at the very first layer.
    • f-R2D (Frame-based 2D ResNet): Applies standard 2D convolutions independently to each of the LL frames with shared weights. No temporal modeling is performed in any convolutional layer; temporal information is merged only at the top via global average pooling.
    • R3D (3D ResNet): Uses full 3D convolutional filters of size Ni1×t×d×dN_{i-1} \times t \times d \times d convolving jointly across time and space throughout all residual stages.
    • MCx (Mixed Convolutions): Employs 3D convolutions in early layers up to convolutional group xx, then switches to 2D spatial convolutions for all deeper groups (groups xx through 5), resting on the hypothesis that motion modeling is a low/mid-level operation.
    • rMCx (Reversed Mixed Convolutions): Employs 2D convolutions in early groups (capturing appearance) and switches to 3D convolutions from group xx onwards to capture higher-level temporal patterns.
    • R(2+1)D: Replaces every 3D convolution homogeneously throughout all layers with a factorized (2+1)D(2+1)\text{D} block (2D spatial convolution followed by 1D temporal convolution).
  4. Knowl 4 — Spatiotemporal ResNet Architecture Specifications and Striding

    model/method

    Video ResNet architectures (R3D and R(2+1)D) use non-bottleneck residual blocks consisting of two convolutional layers with ReLU activations. Given an input clip of size 3×L×112×1123 \times L \times 112 \times 112:

    • conv1: Filter size 3×7×73 \times 7 \times 7, 64 channels, spatial stride 1×2×21 \times 2 \times 2, yielding output size L×56×56L \times 56 \times 56.
    • conv2_x: Residual blocks with 3×3×33 \times 3 \times 3 convolutions, 64 channels (repeated 2 times in 18-layer nets; 3 times in 34-layer nets), output size L×56×56L \times 56 \times 56.
    • conv3_x: Residual blocks with 3×3×33 \times 3 \times 3 convolutions, 128 channels (repeated 2 times in 18-layer nets; 4 times in 34-layer nets), spatiotemporal stride 2×2×22 \times 2 \times 2 applied at conv3_1, output size L2×28×28\frac{L}{2} \times 28 \times 28.
    • conv4_x: Residual blocks with 3×3×33 \times 3 \times 3 convolutions, 256 channels (repeated 2 times in 18-layer nets; 6 times in 34-layer nets), spatiotemporal stride 2×2×22 \times 2 \times 2 applied at conv4_1, output size L4×14×14\frac{L}{4} \times 14 \times 14.
    • conv5_x: Residual blocks with 3×3×33 \times 3 \times 3 convolutions, 512 channels (repeated 2 times in 18-layer nets; 3 times in 34-layer nets), spatiotemporal stride 2×2×22 \times 2 \times 2 applied at conv5_1, output size L8×7×7\frac{L}{8} \times 7 \times 7.

    The final convolutional tensor of size 512×L8×7×7512 \times \frac{L}{8} \times 7 \times 7 is processed by global spatiotemporal average pooling (1×1×11 \times 1 \times 1) to produce a 512-dimensional feature vector, followed by a fully connected classification layer with softmax.

    In R(2+1)D, each 3×3×33 \times 3 \times 3 convolution is replaced with its corresponding spatial (1×3×31 \times 3 \times 3) and temporal (3×1×13 \times 1 \times 1) sequence.

  5. Knowl 5 — Expressiveness and Optimization Benefits of (2+1)D Factorization

    empirical result

    Factorizing 3D convolutions into (2+1)D(2+1)\text{D} blocks provides two distinct advantages over full 3D convolutions of equal parameter budget:

    1. Doubled Non-linearities: Introducing an additional Batch Normalization and ReLU activation between the spatial 2D and temporal 1D convolutions doubles the number of non-linear transformations in the network without increasing parameters, expanding the complexity of functions the network can represent.
    2. Optimization Ease: Explicitly separating spatial and temporal dimensions facilitates gradient-based optimization. When comparing R(2+1)D and R3D models of identical depths (18 and 34 layers) on Kinetics, R(2+1)D achieves consistently lower training error alongside lower validation error across training epochs. The reduction in training error is substantially more pronounced in deeper networks (34 layers), indicating that the optimization advantage increases with network depth.
  6. Knowl 6 — Kinetics-400 Performance Across Spatiotemporal Convolution Types

    data/table

    Evaluating 18-layer ResNet variants on the Kinetics validation set demonstrates that R(2+1)D outperforms all 2D, mixed, and full 3D convolutional models across both 8-frame and 16-frame input settings:

    Net # params Input: 8×112×1128 \times 112 \times 112 Input: 16×112×11216 \times 112 \times 112
    Clip@1 (%) Video@1 (%) Clip@1 (%) Video@1 (%)
    R2D 11.4M 46.7 59.5 47.0 58.9
    f-R2D 11.4M 48.1 59.4 50.3 60.5
    R3D 33.4M 49.4 61.8 52.5 64.2
    MC2 11.4M 50.2 62.5 53.1 64.2
    MC3 11.7M 50.7 62.9 53.7 64.7
    MC4 12.7M 50.5 62.5 53.7 65.1
    MC5 16.9M 50.3 62.5 53.7 65.1
    rMC2 33.3M 49.8 62.1 53.1 64.9
    rMC3 33.0M 49.8 62.3 53.2 65.0
    rMC4 32.0M 49.9 62.3 53.4 65.1
    rMC5 27.9M 49.4 61.2 52.1 63.1
    R(2+1)D 33.3M 52.8 64.8 56.8 68.0

    R(2+1)D attains the highest accuracy on both 8-frame clips (52.8% clip / 64.8% video) and 16-frame clips (56.8% clip / 68.0% video), providing a 3.8% video top-1 accuracy gain over standard R3D at essentially identical parameter count and FLOP complexity. Performance gaps between 2D baselines and spatiotemporal models widen on 16-frame clips, demonstrating the increasing importance of temporal modeling on longer sequences.

  7. Knowl 7 — Clip Length Effects, Inference Aggregation, and Progressive Staged Training

    empirical result

    Experiments using an 18-layer R(2+1)D model on Kinetics reveal key behaviors regarding temporal clip length and training strategies:

    1. Clip Length Scaling: Training on longer clips (from 8 up to 48 frames) monotonically increases clip top-1 accuracy (from 52.8% at 8 frames to over 60% at 48 frames), but video top-1 accuracy peaks at 32 frames (69.4%) and plateaus thereafter.
    2. Training vs. Testing Clip Length Mismatch: Applying a model trained on 8-frame clips directly to 32-frame test clips drops clip accuracy by 1.2% (from 52.8% to 51.6%) and video accuracy by 5.8% (from 64.8% to 59.0%), proving that longer temporal reasoning cannot be attained by simply expanding test-time temporal support without training on longer clips.
    3. Progressive Fine-Tuning Speedup: Training an 18-layer R(2+1)D on 8-frame clips (11.8 hours on 64 GPUs) followed by fine-tuning on 32-frame clips (8.7 hours, total 20.5 hours) yields 59.8% Clip@1 and 68.0% Video@1 accuracy. This achieves performance comparable to training from scratch on 32-frame clips (60.1% Clip@1, 69.4% Video@1 in 59.8 hours) at a fraction of the computational training time.
    4. Inference Clip Sampling: For video-level prediction, averaging predictions across 20 uniformly sampled clips per video is only 0.5% lower in top-1 accuracy compared to averaging 100 clips, while offering a 5×5\times speedup at test time.
  8. Knowl 8 — Sports-1M Benchmark Action Recognition Results

    data/table

    Performance of the 34-layer R(2+1)D model compared against prior methods on the Sports-1M dataset:

    Method Clip@1 (%) Video@1 (%) Video@5 (%)
    DeepVideo 41.9 60.9 80.2
    C3D 46.1 61.1 85.2
    2D ResNet-152 46.5 64.6 86.4
    Conv pooling 71.7 90.4
    P3D 47.9 66.4 87.4
    R3D-RGB-8frame 53.8
    R(2+1)D-RGB-8frame 56.1 72.0 91.2
    R(2+1)D-Flow-8frame 44.5 65.5 87.2
    R(2+1)D-Two-Stream-8frame 72.2 91.4
    R(2+1)D-RGB-32frame 57.0 73.0 91.5
    R(2+1)D-Flow-32frame 46.4 68.4 88.7
    R(2+1)D-Two-Stream-32frame 73.3 91.9

    R(2+1)D-34 RGB (32 frames) achieves 57.0% Clip@1 accuracy, outperforming 152-layer P3D (47.9%) by 9.1%, 2D ResNet-152 (46.5%) by 10.5%, and C3D (46.1%) by 10.9%. The R(2+1)D Two-Stream model (32 frames) reaches 73.3% Video@1 and 91.9% Video@5 accuracy.

  9. Knowl 9 — Kinetics-400 Action Recognition Results and Comparison with I3D

    data/table

    Performance comparison between 34-layer R(2+1)D and I3D architectures on the Kinetics validation set under different modalities and pretraining regimes:

    Method Pretraining Dataset Top-1 (%) Top-5 (%)
    I3D-RGB None 67.5 87.2
    I3D-RGB ImageNet 72.1 90.3
    I3D-Flow ImageNet 65.3 86.2
    I3D-Two-Stream ImageNet 75.7 92.0
    R(2+1)D-RGB None 72.0 90.0
    R(2+1)D-Flow None 67.5 87.2
    R(2+1)D-Two-Stream None 73.9 90.9
    R(2+1)D-RGB Sports-1M 74.3 91.4
    R(2+1)D-Flow Sports-1M 68.5 88.1
    R(2+1)D-Two-Stream Sports-1M 75.4 91.9

    When trained entirely from scratch without pretraining, R(2+1)D-34 RGB achieves 72.0% Top-1 accuracy, exceeding I3D trained from scratch (67.5%) by 4.5%. Pretrained on Sports-1M, R(2+1)D-34 reaches 74.3% (RGB) and 75.4% (Two-Stream).

  10. Knowl 10 — Transfer Learning Performance on UCF101 and HMDB51

    data/table

    Fine-tuning 34-layer R(2+1)D models pretrained on Sports-1M or Kinetics across all 3 splits of UCF101 and HMDB51:

    Method Pretraining Dataset UCF101 (%) HMDB51 (%)
    Two-Stream ImageNet 88.0 59.4
    Action Transform ImageNet 92.4 62.0
    Conv Pooling Sports-1M 88.6
    FST CN ImageNet 88.1 59.1
    Two-Stream Fusion ImageNet 92.5 65.4
    Spatiotemp. ResNet ImageNet 93.4 66.4
    Temporal Segment Net ImageNet 94.2 69.4
    P3D ImageNet + Sports-1M 88.6
    I3D-RGB ImageNet + Kinetics 95.6 74.8
    I3D-Flow ImageNet + Kinetics 96.7 77.1
    I3D-Two-Stream ImageNet + Kinetics 98.0 80.7
    R(2+1)D-RGB Sports-1M 93.6 66.6
    R(2+1)D-Flow Sports-1M 93.3 70.1
    R(2+1)D-TwoStream Sports-1M 95.0 72.7
    R(2+1)D-RGB Kinetics 96.8 74.5
    R(2+1)D-Flow Kinetics 95.5 76.4
    R(2+1)D-TwoStream Kinetics 97.3 78.7

    Kinetics pretraining is more effective than Sports-1M pretraining for transferring to UCF101 and HMDB51. R(2+1)D with Kinetics pretraining achieves 96.8% (RGB) and 97.3% (Two-Stream) on UCF101, and 74.5% (RGB) and 78.7% (Two-Stream) on HMDB51.

Coverage note — None was omitted; all core contributions, models, parameter matching formulations, empirical comparisons, optimization insights, and benchmark evaluations across four datasets are captured.

References

  1. 1.M. Baccouche, F. Mamalet, C. Wolf, C. Garcia, and A. Baskurt. Sequential Deep Learning for Human Action Recognition, pages 29–39. Springer Berlin Heidelberg, Berlin, Heidelberg, 2011. 2
  2. 2.N. Ballas, L. Yao, C. Pal, and A. Courville. Delving deeper into convolutional networks for learning video representations. arXiv preprint arXiv:1511.06432, 2015. 2
  3. 3.Caffe2-Team. Caffe2: A new lightweight, modular, and scalable deep learning framework. https://caffe2.ai/. 6
  4. 4.J. Carreira and A. Zisserman. Quo vadis, action recognition? a new model and the kinetics dataset. In CVPR, 2017. 1, 3, 5, 7, 8
  5. 5.N. Dalal, B. Triggs, and C. Schmid. Human Detection Using Oriented Histograms of Flow and Appearance. In A. Leonardis, H. Bischof, and A. Pinz, editors, European Conference on Computer Vision (ECCV ’06), volume 3952 of Lecture Notes in Computer Science (LNCS), pages 428–441, Graz, Austria, May 2006. Springer-Verlag. 2
  6. 6.P. Dollar, V. Rabaud, G. Cottrell, and S. Belongie. Behavior recognition via sparse spatio-temporal features. In Proc. ICCV VS-PETS, 2005. 2
  7. 7.J. Donahue, L. Anne Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrell. Long-term recurrent convolutional networks for visual recognition and description. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2625–2634, 2015. 2
  8. 8.G. Farneback. Two-frame motion estimation based on polynomial expansion. In Image Analysis, 13th Scandinavian Conference, SCIA 2003, Halmstad, Sweden, June 29 - July 2, 2003, Proceedings, pages 363–370, 2003. 7
  9. 9.C. Feichtenhofer, A. Pinz, and R. P. Wildes. Spatiotemporal residual networks for video action recognition. In NIPS, 2016. 2, 7, 8
  10. 10.C. Feichtenhofer, A. Pinz, and A. Zisserman. Convolutional two-stream network fusion for video action recognition. In CVPR, 2016. 2, 8
  11. 11.R. Girdhar, D. Ramanan, A. Gupta, J. Sivic, and B. C. Russell. Actionvlad: Learning spatio-temporal aggregation for action classification. In CVPR, 2017. 2
  12. 12.P. Goyal, P. Dollar, R. Girshick, P. Noordhuis, L. Wesolowski, A. Kyrola, A. Tulloch, Y. Jia, and K. He. Accurate, large minibatch sgd: training imagenet in 1 hour. arXiv preprint arXiv:1706.02677, 2017. 5
  13. 13.K. He, X. Zhang, S. Ren, and J. Sun. Deep residual learning for image recognition. In CVPR, 2016. 1, 2, 3, 5, 7
  14. 14.G. Huang, Z. Liu, and K. Q. Weinberger. Densely connected convolutional networks. In CVPR, 2017. 1
  15. 15.S. Ji, W. Xu, M. Yang, and K. Yu. 3d convolutional neural networks for human action recognition. IEEE TPAMI, 35(1):221–231, 2013. 1, 2, 3
  16. 16.A. Karpathy, G. Toderici, S. Shetty, T. Leung, R. Sukthankar, and L. Fei-Fei. Large-scale video classification with convolutional neural networks. In CVPR, 2014. 1, 2, 5, 7
  17. 17.W. Kay, J. Carreira, K. Simonyan, B. Zhang, C. Hillier, S. Vijayanarasimhan, F. Viola, T. Green, T. Back, P. Natsev, M. Suleyman, and A. Zisserman. The kinetics human action video dataset. CoRR, abs/1705.06950, 2017. 1
  18. 18.A. Klaser, M. Marszałek, and C. Schmid. A spatio-temporal descriptor based on 3d-gradients. In BMVC, 2008. 2
  19. 19.A. Krizhevsky, I. Sutskever, and G. Hinton. Imagenet classification with deep convolutional neural networks. In NIPS, 2012. 1, 2
  20. 20.H. Kuehne, H. Jhuang, E. Garrote, T. Poggio, and T. Serre. HMDB: a large video database for human motion recognition. In ICCV, 2011. 5, 8
  21. 21.I. Laptev and T. Lindeberg. Space-time interest points. In ICCV, 2003. 2
  22. 22.Q. V. Le, W. Y. Zou, S. Y. Yeung, and A. Y. Ng. Learning hierarchical invariant spatio-temporal features for action recognition with independent subspace analysis. In CVPR, 2011. 2
  23. 23.P. Molchanov, X. Yang, S. Gupta, K. Kim, S. Tyree, and J. Kautz. Online detection and classification of dynamic hand gestures with recurrent 3d convolutional neural network. In CVPR, 2016. 2
  24. 24.Y. Pan, T. Mei, T. Yao, H. Li, and Y. Rui. Jointly modeling embedding and translation to bridge video and language. In CVPR, 2016. 2
  25. 25.Z. Qiu, T. Yao, , and T. Mei. Learning spatio-temporal representation with pseudo-3d residual networks. In ICCV, 2017. 1, 2, 4, 7, 8
  26. 26.S. Sadanand and J. Corso. Action bank: A high-level representation of activity in video. In CVPR, 2012. 2
  27. 27.P. Scovanner, S. Ali, and M. Shah. A 3-dimensional sift descriptor and its application to action recognition. In ACM MM, 2007. 2
  28. 28.Z. Shou, D. Wang, and S.-F. Chang. Temporal action localization in untrimmed videos via multi-stage cnns. In CVPR, 2016. 2
  29. 29.K. Simonyan and A. Zisserman. Two-stream convolutional networks for action recognition in videos. In NIPS, 2014. 2, 3, 7, 8
  30. 30.K. Simonyan and A. Zisserman. Very deep convolutional networks for large-scale image recognition. In ICLR, 2015. 1, 4
  31. 31.K. Soomro, A. R. Zamir, and M. Shah. UCF101: A dataset of 101 human action classes from videos in the wild. In CRCV-TR-12-01, 2012. 5, 8
  32. 32.N. Srivastava, E. Mansimov, and R. Salakhudinov. Unsupervised learning of video representations using lstms. In International Conference on Machine Learning, pages 843–852, 2015. 2
  33. 33.L. Sun, K. Jia, D.-Y. Yeung, and B. E. Shi. Human action recognition using factorized spatio-temporal convolutional networks. In ICCV, 2015. 2, 8
  34. 34.C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich. Going deeper with convolutions. In CVPR, 2015. 1
  35. 35.G. W. Taylor, R. Fergus, Y. LeCun, and C. Bregler. Convolutional learning of spatio-temporal features. In ECCV, pages 140–153. Springer, 2010. 1, 2
  36. 36.D. Tran, L. Bourdev, R. Fergus, L. Torresani, and M. Paluri. Learning spatiotemporal features with 3d convolutional networks. In ICCV, 2015. 1, 2, 3, 7
  37. 37.G. Varol, I. Laptev, and C. Schmid. Long-term Temporal Convolutions for Action Recognition. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2017. 6, 7
  38. 38.H. Wang and C. Schmid. Action recognition with improved trajectories. In ICCV, 2013. 1, 2
  39. 39.L. Wang, Y. Xiong, Z. Wang, Y. Qiao, D. Lin, X. Tang, and L. V. Gool. Temporal segment networks: Towards good practices for deep action recognition. In ECCV, 2016. 2, 8
  40. 40.X. Wang, A. Farhadi, and A. Gupta. Actions ˜ transformations. In CVPR, 2016. 2, 8
  41. 41.Z. Xu, Y. Yang, and A. G. Hauptmann. A discriminative cnn video representation for event detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 1798–1807, 2015. 2
  42. 42.J. Yue-Hei Ng, M. Hausknecht, S. Vijayanarasimhan, O. Vinyals, R. Monga, and G. Toderici. Beyond short snippets: Deep networks for video classification. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 4694–4702, 2015. 2, 7, 8
  43. 43.C. Zach, T. Pock, and H. Bischof. A duality based approach for realtime tv-l 1 optical flow. Pattern Recognition, pages 214–223, 2007. 8

Citation

MLA
Tran, D., et al. “A Closer Look at Spatiotemporal Convolutions for Action Recognition”. arXiv, 2017, http://arxiv.org/abs/1711.11248v3.
APA
Tran, D., Wang, H., Torresani, L., Ray, J., LeCun, Y., & Paluri, M. (2017). A Closer Look at Spatiotemporal Convolutions for Action Recognition. arXiv. http://arxiv.org/abs/1711.11248v3
Chicago
Tran, D., H. Wang, L. Torresani, J. Ray, Y. LeCun, and M. Paluri. 2017. “A Closer Look at Spatiotemporal Convolutions for Action Recognition”. arXiv. http://arxiv.org/abs/1711.11248v3.
Harvard
Tran, D. et al. (2017) “A Closer Look at Spatiotemporal Convolutions for Action Recognition”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1711.11248v3.
Vancouver
1. Tran D, Wang H, Torresani L, Ray J, LeCun Y, Paluri M (2017) A Closer Look at Spatiotemporal Convolutions for Action Recognition. arXiv

BibTeX

@article{tran2017closer,
  title = {A Closer Look at Spatiotemporal Convolutions for Action Recognition},
  author = {Tran, Du and Wang, Heng and Torresani, Lorenzo and Ray, Jamie and LeCun, Yann and Paluri, Manohar},
  year = {2017},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1711.11248v3},
  eprint = {1711.11248}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF

License: IEEE