Can Spatiotemporal 3D CNNs Retrace the History of 2D CNNs and ImageNet?

Kensho HaraHirokatsu KataokaYutaka Satoh

article2017CVPR2,271 citationsAIST Best Paper

Demonstrates that the Kinetics dataset enables successful training of deep 3D CNN architectures up to 152 layers without overfitting, providing effective pretrained video models that outperform complex 2D approaches on action recognition benchmarks.

Listen

Recent advances in image recognition have relied heavily on training very deep two-dimensional neural networks on massive datasets like ImageNet, which in turn provided powerful baseline models for many visual applications. In video analysis, however, progress has lagged because three-dimensional networkswhich process spatial and temporal motion data simultaneouslyrequire massive amounts of data to avoid severe memorization issues, known as overfitting. Historically, standard video benchmark collections were far too small to support deep architectures, leaving researchers uncertain whether newer, larger video collections could enable the same breakthroughs seen in still-image processing.

The article evaluates whether modern video datasets provide sufficient scale to train very deep three-dimensional convolutional neural networks directly from scratch and whether models pre-trained on these large video collections can successfully transfer visual knowledge to smaller video tasks.

To address this question, the researchers conducted extensive empirical experiments using four benchmark video datasets of varying sizes: UCF-101, HMDB-51, ActivityNet, and Kinetics, the last comprising over 300,000 video clips across 400 action categories. They systematically evaluated a range of residual network architectures, varying model depths from shallow 18-layer configurations up to very deep 200-layer models, alongside advanced structural variants such as Wide ResNet, ResNeXt, and DenseNet. The models were trained from scratch to evaluate convergence and overfitting, and the top-performing models were subsequently fine-tuned on smaller target benchmarks to measure transfer learning effectiveness.

The experiments produced four critical findings. First, training even the shallowest 18-layer models from scratch on UCF-101, HMDB-51, and ActivityNet resulted in severe overfitting and poor validation accuracy (ranging from 16.2% to 40.1%), proving these collections remain inadequate for training three-dimensional networks from scratch. Second, the Kinetics dataset proved large enough to train models up to 152 layers deep without overfitting, achieving continuous accuracy gains that closely parallel the historical trajectory of 2D models on ImageNet. Third, among the tested architectures, ResNeXt-101 demonstrated superior performance, reaching a 78.4% average top-1/top-5 accuracy on the Kinetics test set when processing longer 64-frame input clips. Fourth, pre-training deep models on Kinetics and fine-tuning them on smaller datasets delivered outstanding performance, achieving 94.5% accuracy on UCF-101 and 70.2% on HMDB-51, outperforming complex, traditional two-dimensional video processing pipelines.

These results demonstrate that large-scale video pre-training is viable and effective, establishing Kinetics as a functional equivalent of ImageNet for video domains. For technical leaders and developers, this means that simpler, unified three-dimensional architectures pre-trained on large video corpuses can replace complex multi-pipeline systems, reducing engineering overhead while improving operational performance and accuracy.

Organizations developing video understanding systems should adopt pre-trained deep three-dimensional modelsspecifically high-capacity architectures like ResNeXt-101as their standard foundation, fine-tuning only the final classification layers for specific target tasks. Future development should focus on extending these pre-trained representations to broader operational tasks beyond classification, including action detection, automated video summarization, and motion estimation.

Confidence in these conclusions is high regarding action classification benchmarks, as the findings are validated across multiple standard datasets and network depths. However, decision-makers should recognize boundary conditions: the findings rely heavily on trimmed, curated short video clips, and high-performing deep three-dimensional models require substantial graphics processing capacity and memory during both training and inference.

Cover for Can Spatiotemporal 3D CNNs Retrace the History of 2D CNNs and ImageNet?

Abstract

The purpose of this study is to determine whether current video datasets have sufficient data for training very deep convolutional neural networks (CNNs) with spatio-temporal three-dimensional (3D) kernels. Recently, the performance levels of 3D CNNs in the field of action recognition have improved significantly. However, to date, conventional research has only explored relatively shallow 3D architectures. We examine the architectures of various 3D CNNs from relatively shallow to very deep ones on current video datasets. Based on the results of those experiments, the following conclusions could be obtained: (i) ResNet-18 training resulted in significant overfitting for UCF-101, HMDB-51, and ActivityNet but not for Kinetics. (ii) The Kinetics dataset has sufficient data for training of deep 3D CNNs, and enables training of up to 152 ResNets layers, interestingly similar to 2D ResNets on ImageNet. ResNeXt-101 achieved 78.4% average accuracy on the Kinetics test set. (iii) Kinetics pretrained simple 3D architectures outperforms complex 2D architectures, and the pretrained ResNeXt-101 achieved 94.5% and 70.2% on UCF-101 and HMDB-51, respectively. The use of 2D CNNs trained on ImageNet has produced significant progress in various tasks in image. We believe that using deep 3D CNNs together with Kinetics will retrace the successful history of 2D CNNs and ImageNet, and stimulate advances in computer vision for videos. The codes and pretrained models used in this study are publicly available. this https URL

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Video Datasets
  • 2.2 Action Recognition Approaches
  • 3 Experimental configuration
  • 3.1 Summary
  • 3.2 Network architectures
  • 3.3 Implementation
  • 3.4 Datasets
  • 4 Results and discussion
  • 4.1 Analyses of training on each dataset
  • 4.2 Analyses of deeper networks
  • 4.3 Analyses of fine-tuning
  • 5 Conclusion
  • References

Knowls

  1. Knowl 1 — Architectures of 3D Residual, ResNeXt, Wide ResNet, and DenseNet Models for Video Recognition

    model/method

    Spatiotemporal 3D convolutional neural networks extend 2D deep architectures to video inputs of shape 3×T×H×W3 \times T \times H \times W (channels ×\times frames ×\times height ×\times width). Across all models, the initial layer (conv1) is a 7×7×77 \times 7 \times 7 3D convolution with 64 filters, temporal stride 1, and spatial stride 2, followed by batch normalization (BN), ReLU, and a 3×3×33 \times 3 \times 3 max-pooling layer with stride 2.

    Spatio-temporal downsampling is subsequently executed with a stride of 2 at the beginning of stages conv3_1, conv4_1, and conv5_1 (except in DenseNet).

    1. 3D ResNet (Basic Block - Depths 18, 34): Blocks consist of two 3×3×33 \times 3 \times 3 convolutional layers, each followed by BN and ReLU. Shortcuts utilize identity mappings and zero-padding for dimension matching.

    2. 3D ResNet (Bottleneck Block - Depths 50, 101, 152, 200): Blocks contain a 1×1×11 \times 1 \times 1 conv (FF channels), a 3×3×33 \times 3 \times 3 conv (FF channels), and a 1×1×11 \times 1 \times 1 conv (4F4F channels), with BN and ReLU after each convolution. Projection shortcuts are used when increasing dimensions.

    3. 3D Pre-activation ResNet-200: Adopts the bottleneck structure with pre-activation ordering (BN \to ReLU \to Conv).

    4. 3D Wide ResNet-50 (WRN-50): Preserves the 50-layer bottleneck topology while applying a widening factor of 2, doubling channel counts to F{128,256,512,1024}F \in \{128, 256, 512, 1024\} across stages conv2_x through conv5_x.

    5. 3D ResNeXt-101: Modifies the bottleneck 3×3×33 \times 3 \times 3 convolution into a grouped convolution with a cardinality of 32 groups, using F{128,256,512,1024}F \in \{128, 256, 512, 1024\}.

    6. 3D DenseNet (Depths 121, 201): Implements pre-activation dense connectivity with a growth rate of 32. Transition layers between stages consist of a 3×3×33 \times 3 \times 3 convolution followed by a 2×2×22 \times 2 \times 2 average pooling layer with stride 2.

    All architectures conclude with global spatiotemporal average pooling and a fully-connected layer projecting to CC class logits followed by a softmax activation.

  2. Knowl 2 — Data Sampling, Multi-Scale Augmentation, and Inference Protocol for 3D CNNs

    experimental setup

    The training and evaluation protocols for 3D spatio-temporal video CNNs operate as follows:

    Training Sample Generation & Augmentation:

    • A temporal position is uniformly sampled in the video, and a 16-frame clip is extracted around it. Videos shorter than 16 frames are looped.
    • A spatial crop location is randomly chosen from the four corners or the center.
    • Multi-scale spatial cropping is applied by selecting a crop scale s{1,1/21/4,1/2,1/23/4,1/2}s \in \{1, 1/2^{1/4}, 1/\sqrt{2}, 1/2^{3/4}, 1/2\}, where s=1s=1 corresponds to the shorter frame edge length and s=0.5s=0.5 corresponds to half that length. The cropped region has an aspect ratio of 1.
    • The sample is resized to 112×112112 \times 112 pixels, resulting in an input volume of 3×16×112×1123 \times 16 \times 112 \times 112.
    • The volume is horizontally flipped with a 50% probability, and per-channel mean subtraction using ActivityNet channel means is applied.

    Optimization:

    • Models are trained using mini-batch stochastic gradient descent (SGD) with momentum 0.9, cross-entropy loss, and weight decay 10310^{-3} (scratch) or 10510^{-5} (fine-tuning).
    • Training from scratch begins with an initial learning rate of 0.1, divided by 10 whenever the validation loss plateaus.

    Inference (Video-Level Prediction):

    • Each input video is partitioned into non-overlapping 16-frame clips using a sliding window.
    • Each clip is spatially cropped at the center at scale s=1s=1 and evaluated by the model.
    • The predicted class probability vectors are averaged across all clips of the video, and the argmax determines the video label.
  3. Knowl 3 — Overfitting on Small Video Benchmarks vs Scratch Training on Kinetics

    empirical result

    When training a shallow 3D CNN (ResNet-18) from scratch on smaller video datasets, severe overfitting occurs:

    • On UCF-101 (split 1, ~13.3K videos), HMDB-51 (split 1, ~6.8K videos), and ActivityNet v1.3 (~28.1K activity instances, 849 hours), validation loss converges quickly to high values while training loss continues to drop.
    • Per-clip validation top-1 accuracies are 40.1% on UCF-101, 16.2% on HMDB-51, and 26.8% on ActivityNet.
    • Video-level top-1 classification accuracies from scratch reach only 42.4% on UCF-101 and 17.1% on HMDB-51.

    Conversely, when trained on the Kinetics dataset (~300K trimmed videos across 400 action classes; ~240K train, ~20K validation), the validation loss remains close to the training loss throughout optimization without suffering from overfitting. This demonstrates that Kinetics provides sufficient data volume to optimize deep 3D CNN parameters from scratch.

  4. Knowl 4 — Depth Scaling and Architecture Comparison on the Kinetics Validation Benchmark

    data/table

    Top-1, Top-5, and average (mean of Top-1 and Top-5) classification accuracies on the Kinetics validation set (~20,000 videos across 400 categories) for various 3D architectures trained from scratch on 16-frame 112×112112 \times 112 clips:

    Method Top-1 (%) Top-5 (%) Average (%)
    ResNet-18 54.2 78.1 66.1
    ResNet-34 60.1 81.9 71.0
    ResNet-50 61.3 83.1 72.2
    ResNet-101 62.8 83.9 73.3
    ResNet-152 63.0 84.4 73.7
    ResNet-200 63.1 84.4 73.7
    ResNet-200 (pre-act) 63.0 83.7 73.4
    Wide ResNet-50 64.1 85.3 74.7
    ResNeXt-101 65.1 85.7 75.4
    DenseNet-121 59.7 81.9 70.8
    DenseNet-201 61.3 83.3 72.3

    Performance increases consistently with depth up to 152 layers, saturating at 200 layers. Wide ResNet-50 outperforms standard ResNet-152, and ResNeXt-101 achieves the highest overall accuracy (75.4% average), demonstrating that cardinality improvements translate effectively to 3D convolutional networks.

  5. Knowl 5 — Kinetics Test Set Evaluation of 3D ResNeXt Against State-of-the-Art Approaches

    data/table

    Performance of models trained from scratch on the Kinetics test set (~40,000 videos, 400 classes). The evaluation contrasts 3D ResNeXt-101 with standard 16-frame inputs and extended 64-frame inputs (ResNeXt-101 (64f), input shape 3×64×112×1123 \times 64 \times 112 \times 112) against 2D and 3D action recognition baselines:

    Method Top-1 (%) Top-5 (%) Average (%)
    ResNeXt-101 (16 frames) 74.5
    ResNeXt-101 (64 frames) 78.4
    CNN+LSTM 57.0 79.0 68.0
    Two-stream CNN 61.0 81.3 71.2
    C3D w/ BN 56.1 79.5 67.8
    RGB-I3D 68.4 88.0 78.2
    Two-stream I3D 71.6 90.0 80.8

    ResNeXt-101 with 64-frame input clips achieves 78.4% average accuracy, outperforming the scratch-trained RGB-I3D baseline (78.2%) despite having an input spatial resolution four times smaller (112×112112 \times 112 vs. 224×224224 \times 224).

  6. Knowl 6 — Partial-Layer Fine-Tuning Strategy for Kinetics-Pretrained 3D CNNs

    model/method

    To transfer representations learned on Kinetics to smaller action recognition datasets (such as UCF-101 and HMDB-51), weights are initialized from models pretrained from scratch on Kinetics.

    All early feature extraction layers up to stage conv4_x are frozen, and gradient updates are restricted strictly to the final residual stage conv5_x and the linear classification head. Fine-tuning uses SGD with an initial learning rate of 10310^{-3} and weight decay 10510^{-5}, maintaining identical parameter counts across depths 50 through 200 during target adaptation.

  7. Knowl 7 — Transfer Learning Performance on UCF-101 and HMDB-51 Using Kinetics Pretraining

    data/table

    Top-1 action classification accuracy (averaged across the three standard train/test splits) on UCF-101 and HMDB-51 using 3D architectures pretrained on Kinetics-400:

    Method UCF-101 Top-1 (%) HMDB-51 Top-1 (%)
    ResNet-18 (scratch) 42.4 17.1
    ResNet-18 (pretrained) 84.4 56.4
    ResNet-34 (pretrained) 87.7 59.1
    ResNet-50 (pretrained) 89.3 61.0
    ResNet-101 (pretrained) 88.9 61.7
    ResNet-152 (pretrained) 89.6 62.4
    ResNet-200 (pretrained) 89.6 63.5
    DenseNet-121 (pretrained) 87.6 59.6
    ResNeXt-101 (pretrained) 90.7 63.8

    Pretraining on Kinetics provides substantial accuracy gains over training from scratch (+42.0% on UCF-101 and +39.3% on HMDB-51 for ResNet-18). Performance scales with pretrained depth, with ResNeXt-101 achieving 90.7% on UCF-101 and 63.8% on HMDB-51.

  8. Knowl 8 — Comparison of Kinetics-Pretrained 3D ResNeXt Against State-of-the-Art 2D and 3D Baselines on UCF-101 and HMDB-51

    data/table

    Comparison of top-1 action classification accuracy (averaged over 3 splits) on UCF-101 and HMDB-51 benchmarks across 2D and 3D methods:

    Method Dimension UCF-101 Top-1 (%) HMDB-51 Top-1 (%)
    ResNeXt-101 (16 frames) 3D 90.7 63.8
    ResNeXt-101 (64 frames) 3D 94.5 70.2
    C3D 3D 82.3
    P3D 3D 88.6
    Two-stream I3D 3D 98.0 80.7
    Two-stream CNN 2D 88.0 59.4
    TDD 2D 90.3 63.2
    ST Multiplier Net 2D 94.2 68.9
    TSN 2D 94.2 69.4

    When using 64-frame input clips, Kinetics-pretrained single-stream RGB 3D ResNeXt-101 achieves 94.5% on UCF-101 and 70.2% on HMDB-51, surpassing complex multi-stream 2D methods such as Spatiotemporal Multiplier Networks (94.2% / 68.9%) and Temporal Segment Networks (94.2% / 69.4%).

Coverage note — None was omitted; all contributed network specifications, training protocols, depth scaling experiments, test benchmarks, and transfer learning evaluations are included.

References

  1. 1.S. Abu-El-Haija, N. Kothari, J. Lee, P. Natsev, G. Toderici, B. Varadarajan, and S. Vijayanarasimhan. YouTube-8M: A large-scale video classification benchmark. arXiv preprint, arXiv:1609.08675, 2016. 3
  2. 2.J. Carreira and A. Zisserman. Quo vadis, action recognition? A new model and the Kinetics dataset. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 4724–4733, 2017. 2, 3
  3. 3.J. Carreira and A. Zisserman. Quo vadis, action recognition? A new model and the Kinetics dataset. arXiv preprint, arXiv:1705.07750, 2017. 7, 8
  4. 4.J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei. ImageNet: A Large-Scale Hierarchical Image Database. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2009. 1
  5. 5.B. G. Fabian Caba Heilbron, Victor Escorcia and J. C. Niebles. ActivityNet: A large-scale video benchmark for human activity understanding. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 961–970, 2015. 1, 2, 3, 5, 6
  6. 6.C. Feichtenhofer, A. Pinz, and R. Wildes. Spatiotemporal residual networks for video action recognition. In Proceedings of the Advances in Neural Information Processing Systems (NIPS), pages 3468–3476, 2016. 3
  7. 7.C. Feichtenhofer, A. Pinz, and R. P. Wildes. Spatiotemporal multiplier networks for video action recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017. 3, 8
  8. 8.C. Feichtenhofer, A. Pinz, and A. Zisserman. Convolutional two-stream network fusion for video action recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 1933–1941, 2016. 3
  9. 9.K. Hara, H. Kataoka, and Y. Satoh. Learning spatio-temporal features with 3D residual networks for action recognition. In Proceedings of the ICCV Workshop on Action, Gesture, and Emotion Recognition, 2017. 2, 3, 4, 6
  10. 10.K. He, X. Zhang, S. Ren, and J. Sun. Deep residual learning for image recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 770–778, 2016. 1, 2, 3, 4, 5
  11. 11.K. He, X. Zhang, S. Ren, and J. Sun. Identity mappings in deep residual networks. In Proceedings of the European Conference on Computer Vision (ECCV), pages 630–645, 2016. 2, 4, 5, 7
  12. 12.G. Huang, Z. Liu, L. van der Maaten, and K. Q. Weinberger. Densely connected convolutional networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 4700–4708, 2017. 2, 4, 5, 7
  13. 13.S. Ioffe and C. Szegedy. Batch normalization: Accelerating deep network training by reducing internal covariate shift. In Proceedings of the International Conference on Machine Learning, pages 448–456, 2015. 4
  14. 14.S. Ji, W. Xu, M. Yang, and K. Yu. 3D convolutional neural networks for human action recognition. IEEE Transactions on Pattern Analysis and Machine Intelligence, 35(1):221–231, 2013. 2
  15. 15.A. Karpathy, G. Toderici, S. Shetty, T. Leung, R. Sukthankar, and L. Fei-Fei. Large-scale video classification with convolutional neural networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 1725–1732, 2014. 3
  16. 16.W. Kay, J. Carreira, K. Simonyan, B. Zhang, C. Hillier, S. Vijayanarasimhan, F. Viola, T. Green, T. Back, P. Natsev, M. Suleyman, and A. Zisserman. The Kinetics human action video dataset. arXiv preprint, arXiv:1705.06950, 2017. 2, 3, 5, 6, 7
  17. 17.H. Kuehne, H. Jhuang, E. Garrote, T. Poggio, and T. Serre. HMDB: a large video database for human motion recognition. In Proceedings of the International Conference on Computer Vision (ICCV), pages 2556–2563, 2011. 1, 2, 3, 5
  18. 18.V. Nair and G. E. Hinton. Rectified linear units improve restricted boltzmann machines. In Proceedings of the International Conference on Machine Learning, pages 807–814. Omnipress, 2010. 4
  19. 19.Z. Qiu, T. Yao, and T. Mei. Learning spatio-temporal representation with pseudo-3d residual networks. In Proceedings of the International Conference on Computer Vision (ICCV), 2017. 8
  20. 20.K. Simonyan and A. Zisserman. Two-stream convolutional networks for action recognition in videos. In Proceedings of the Advances in Neural Information Processing Systems (NIPS), pages 568–576, 2014. 2, 3, 8
  21. 21.K. Soomro, A. Roshan Zamir, and M. Shah. UCF101: A dataset of 101 human action classes from videos in the wild. CRCV-TR-12-01, 2012. 1, 2, 3, 5
  22. 22.C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich. Going deeper with convolutions. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 1–9, 2015. 3
  23. 23.D. Tran, L. Bourdev, R. Fergus, L. Torresani, and M. Paluri. Learning spatiotemporal features with 3D convolutional networks. In Proceedings of the International Conference on Computer Vision (ICCV), pages 4489–4497, 2015. 2, 8
  24. 24.D. Tran, J. Ray, Z. Shou, S. Chang, and M. Paluri. Convnet architecture search for spatiotemporal feature learning. arXiv preprint, arXiv:1708.05038, 2017. 2, 3, 4, 6
  25. 25.G. Varol, I. Laptev, and C. Schmid. Long-term temporal convolutions for action recognition. arXiv preprint, arXiv:1604.04494, 2016. 2, 3
  26. 26.H. Wang, A. Kläser, C. Schmid, and C.-L. Liu. Dense trajectories and motion boundary descriptors for action recognition. International Journal of Computer Vision, 103(1):60–79, 2013. 6
  27. 27.L. Wang, Y. Qiao, and X. Tang. Action recognition with trajectory-pooled deep-convolutional descriptors. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 4305–4314, 2015. 3, 8
  28. 28.L. Wang, Y. Xiong, Z. Wang, and Y. Qiao. Towards good practices for very deep two-stream convnets. arXiv preprint, arXiv:1507.02159, 2015. 3, 5
  29. 29.L. Wang, Y. Xiong, Z. Wang, Y. Qiao, D. Lin, X. Tang, and L. Van Gool. Temporal segment networks: Towards good practices for deep action recognition. In Proceedings of the European Conference on Computer Vision (ECCV), pages 20–36, 2016. 3, 8
  30. 30.S. Xie, R. Girshick, P. Dollár, Z. Tu, and K. He. Aggregated residual transformations for deep neural networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 1492–1500, 2017. 2, 4, 5, 7
  31. 31.S. Zagoruyko and N. Komodakis. Wide residual networks. In Proceedings of the British Machine Vision Conference, 2016. 2, 4, 5, 7

Citation

MLA
Hara, K., et al. “Can Spatiotemporal 3D CNNs Retrace the History of 2D CNNs and ImageNet?”. arXiv, 2017, http://arxiv.org/abs/1711.09577v2.
APA
Hara, K., Kataoka, H., & Satoh, Y. (2017). Can Spatiotemporal 3D CNNs Retrace the History of 2D CNNs and ImageNet?. arXiv. http://arxiv.org/abs/1711.09577v2
Chicago
Hara, K., H. Kataoka, and Y. Satoh. 2017. “Can Spatiotemporal 3D CNNs Retrace the History of 2D CNNs and ImageNet?”. arXiv. http://arxiv.org/abs/1711.09577v2.
Harvard
Hara, K., Kataoka, H. and Satoh, Y. (2017) “Can Spatiotemporal 3D CNNs Retrace the History of 2D CNNs and ImageNet?”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1711.09577v2.
Vancouver
1. Hara K, Kataoka H, Satoh Y (2017) Can Spatiotemporal 3D CNNs Retrace the History of 2D CNNs and ImageNet?. arXiv

BibTeX

@article{hara2017can,
  title = {Can Spatiotemporal 3D CNNs Retrace the History of 2D CNNs and ImageNet?},
  author = {Hara, Kensho and Kataoka, Hirokatsu and Satoh, Yutaka},
  year = {2017},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1711.09577v2},
  eprint = {1711.09577}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF

License: IEEE