Learning Spatio-Temporal Representation with Pseudo-3D Residual Networks

Zhaofan QiuTing YaoTao Mei

article2017ICCV1,839 citations

Proposes Pseudo-3D Residual Networks (P3D ResNet), an architecture that factorizes 3D convolutions into separable 2D spatial and 1D temporal filters within residual blocks, drastically cutting computation and memory costs while outperforming standard 3D CNNs on video classification benchmarks.

Listen

Digital video content has grown rapidly, driving a strong need for automated multimedia analysis in applications such as action and scene recognition. While convolutional neural networks excel at image understanding, extending them directly to full three-dimensional video convolutions requires immense computing memory and processing cost. Conversely, standard two-dimensional frame-based models and recurrent architectures often fail to capture motion connections across low-level features, limiting their ability to model complex temporal dynamics.

The article evaluates whether decomposing three-dimensional convolutions into separate spatial and temporal operations within a residual neural network architecture can achieve state-of-the-art video understanding while keeping computational complexity and model size manageable.

To test this concept, the authors created three modular building blocks combining two-dimensional spatial filters with one-dimensional temporal filters in parallel and cascaded configurations. They then introduced the Pseudo-3D Residual Net (P3D ResNet), an architecture that interleaves these three block designs to increase structural diversity across network layers. The model leverages pre-trained two-dimensional image features from ImageNet and was trained on over one million videos from the Sports-1M dataset before being evaluated across five benchmark datasets covering action recognition, action similarity labeling, and scene recognition.

The experimental findings show significant performance advantages. On the Sports-1M classification benchmark, P3D ResNet achieved a 66.4% top-1 video accuracy, surpassing conventional three-dimensional networks by 5.3% and standard deep two-dimensional networks by 1.8%, while maintaining a compact 261 MB model size. In action recognition benchmarks, P3D ResNet achieved 88.6% accuracy on UCF101 using frame inputs alone—outperforming competing frame-based methods—and reached 93.7% when combined with trajectory features. On ActivityNet, the architecture improved top-1 accuracy to 75.12%, outperforming standard two-dimensional networks by 3.7% and three-dimensional networks by 9.3%. It also achieved top scores on action similarity (80.8% accuracy on ASLAN) and scene classification (99.5% accuracy on YUPENN and 94.6% on Dynamic Scene).

These results indicate that decoupling spatial and temporal convolutions provides an effective, computationally efficient path to deep video representation. Organizations can achieve higher accuracy without the prohibitive hardware costs and long training times typically required to train deep three-dimensional networks from scratch. Furthermore, initializing spatial layers with existing image recognition models helps bridge data domain gaps and preserves performance even when representations are compressed to lower dimensions.

Based on these outcomes, organizations developing video analytics solutions should consider adopting factorized pseudo-three-dimensional architectures instead of standard full three-dimensional convolutions. For future system enhancement, practitioners and researchers should explore incorporating attention mechanisms, training on longer video clips, and integrating additional modalities such as optical flow and audio streams.

Confidence in these findings is high across standard academic benchmarks, though stakeholders should note certain limitations. About 9.2% of the initial Sports-1M dataset was inaccessible during training due to expired links, and tests relied primarily on short 16-frame clips covering less than half a second. Deployments requiring long-range temporal reasoning over minutes or hours may need additional architectural adaptation.

Cover for Learning Spatio-Temporal Representation with Pseudo-3D Residual Networks

Abstract

Convolutional Neural Networks (CNN) have been regarded as a powerful class of models for image recognition problems. Nevertheless, it is not trivial when utilizing a CNN for learning spatio-temporal video representation. A few studies have shown that performing 3D convolutions is a rewarding approach to capture both spatial and temporal dimensions in videos. However, the development of a very deep 3D CNN from scratch results in expensive computational cost and memory demand. A valid question is why not recycle off-the-shelf 2D networks for a 3D CNN. In this paper, we devise multiple variants of bottleneck building blocks in a residual learning framework by simulating 3×3×33\times3\times3 convolutions with 1×3×31\times3\times3 convolutional filters on spatial domain (equivalent to 2D CNN) plus 3×1×13\times1\times1 convolutions to construct temporal connections on adjacent feature maps in time. Furthermore, we propose a new architecture, named Pseudo-3D Residual Net (P3D ResNet), that exploits all the variants of blocks but composes each in different placement of ResNet, following the philosophy that enhancing structural diversity with going deep could improve the power of neural networks. Our P3D ResNet achieves clear improvements on Sports-1M video classification dataset against 3D CNN and frame-based 2D CNN by 5.3% and 1.8%, respectively. We further examine the generalization performance of video representation produced by our pre-trained P3D ResNet on five different benchmarks and three different tasks, demonstrating superior performances over several state-of-the-art techniques.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 P3D Blocks and P3D ResNet
  • 3.1 3D Convolutions
  • 3.2 Pseudo-3D Blocks
  • 3.3 Pseudo-3D ResNet
  • 4 Spatio-Temporal Representation Learning
  • 5 Video Representation Evaluation
  • 6 Conclusion
  • References

Knowls

  1. Knowl 1 — Pseudo-3D Convolutional Factorization

    model/method

    Standard 3D convolutional networks encode spatio-temporal video information using 3D filter kernels of size d×k×kd \times k \times k, where dd is the kernel temporal depth and kk is the kernel spatial extent (e.g., 3×3×33 \times 3 \times 3). Pseudo-3D (P3D) factorizes a 3×3×33 \times 3 \times 3 convolution into a 2D spatial convolution with filter size 1×3×31 \times 3 \times 3 (denoted by SS) and a 1D temporal convolution with filter size 3×1×13 \times 1 \times 1 (denoted by TT).

    This decoupling reduces network parameter count and memory footprint compared to full 3D CNNs. It also allows the 2D spatial convolutional filters (1×3×31 \times 3 \times 3) to be directly initialized from 2D CNN filters (3×33 \times 3) pre-trained on large-scale image datasets (such as ImageNet), while the 1D temporal filters (3×1×13 \times 1 \times 1) are trained to capture frame-to-frame temporal relationships.

  2. Knowl 2 — Pseudo-3D Bottleneck Building Blocks (P3D-A, P3D-B, P3D-C)

    model/method

    Three distinct bottleneck building block designs incorporate 2D spatial filters SS (1×3×31 \times 3 \times 3) and 1D temporal filters TT (3×1×13 \times 1 \times 1) into a residual learning formulation. Let xtx_t and xt+1x_{t+1} denote the input and output feature representations of the tt-th unit, respectively.

    1. P3D-A (Cascaded): Temporal 1D filters follow spatial 2D filters in series, with direct influence between the two filters on the same path:

    xt+1=(I+T⋅S)⋅xt=xt+T(S(xt))x_{t+1} = (I + T \cdot S) \cdot x_t = x_t + T(S(x_t))

    1. P3D-B (Parallel): Spatial and temporal filters operate in separate parallel paths and their outputs are summed before adding to the identity connection:

    xt+1=(I+S+T)⋅xt=xt+S(xt)+T(xt)x_{t+1} = (I + S + T) \cdot x_t = x_t + S(x_t) + T(x_t)

    1. P3D-C (Cascaded with Spatial Shortcut): Temporal 1D filters follow spatial 2D filters in a cascaded path, while an additional direct shortcut connects the spatial filter output S(xt)S(x_t) to the final accumulation:

    xt+1=(I+S+T⋅S)⋅xt=xt+S(xt)+T(S(xt))x_{t+1} = (I + S + T \cdot S) \cdot x_t = x_t + S(x_t) + T(S(x_t))

    To manage computational complexity, all three blocks use a bottleneck design: two 1×1×11 \times 1 \times 1 convolutions are placed at the input and output of the block to reduce and restore feature channel dimensions around the SS and TT filter layers.

  3. Knowl 3 — Pseudo-3D Residual Net Architecture via Interleaved Structural Diversity

    model/method

    Pseudo-3D Residual Net (P3D ResNet) constructs very deep spatio-temporal networks by replacing the residual units in a ResNet backbone (such as ResNet-50 or ResNet-152) with an alternating sequence of P3D bottleneck blocks.

    Rather than using a single homogeneous block type, P3D ResNet interleaves the three block variants in a cyclic order:

    P3D-A⟶P3D-B⟶P3D-C⟶P3D-A⟶…\text{P3D-A} \longrightarrow \text{P3D-B} \longrightarrow \text{P3D-C} \longrightarrow \text{P3D-A} \longrightarrow \dots

    This structural diversity exposes intermediate representations to cascaded, parallel, and multi-scale temporal connections across different depths, enabling spatio-temporal modeling at all network levels from low-level edges to high-level semantics.

  4. Knowl 4 — Performance and Efficiency of P3D ResNet Variants on UCF101

    data/table

    Ablation experiments comparing ResNet-50 with homogeneous and mixed P3D ResNet-50 variants on UCF101 (split 1). Input clips are 16×160×16016 \times 160 \times 160, initialized from ImageNet-pretrained ResNet-50 weights. Inference speed was measured on a single NVIDIA K40 GPU.

    Method Model size Speed Accuracy
    ResNet-50 92MB 15.0 frame/s 80.8%
    P3D-A ResNet 98MB 9.0 clip/s 83.7%
    P3D-B ResNet 98MB 8.8 clip/s 82.8%
    P3D-C ResNet 98MB 8.6 clip/s 83.0%
    P3D ResNet 98MB 8.8 clip/s 84.2%

    All three homogeneous P3D variants outperform the 2D ResNet-50 baseline. Interleaving all three blocks (P3D ResNet) achieves the highest accuracy (84.2%), improving by 0.5%, 1.4%, and 1.2% over P3D-A, P3D-B, and P3D-C variants alone, demonstrating the benefit of structural diversity.

  5. Knowl 5 — Spatio-Temporal Video Classification Performance on Sports-1M

    data/table

    Performance of a 152-layer P3D ResNet trained on the 1.02 million available videos in the Sports-1M dataset compared against 2D CNN, 3D CNN, and pooling baselines. Clip-level predictions are obtained on center crops of 16-frame clips, and 20 clip predictions are averaged for the video-level score.

    Method Pre-train Data Clip Length Video hit@1 Video hit@5
    Deep Video (Single Frame) ImageNet1K 1 59.3% 77.7%
    Deep Video (Slow Fusion) ImageNet1K 10 60.9% 80.2%
    Convolutional Pooling ImageNet1K 120 72.3% 90.8%
    C3D – 16 60.0% 84.4%
    C3D I380K 16 61.1% 85.2%
    ResNet-152 ImageNet1K 1 64.6% 86.4%
    P3D ResNet (ours) ImageNet1K 16 66.4% 87.4%

    P3D ResNet achieves 66.4% top-1 video hit rate, surpassing the 152-layer 2D ResNet-152 baseline by 1.8% and the 11-layer 3D CNN (C3D) by 5.3% while maintaining a smaller model footprint (261MB vs 321MB for C3D).

  6. Knowl 6 — Action Recognition Evaluation on UCF101 and ActivityNet

    data/table

    Evaluation of P3D ResNet video representations (2,048-dimensional features extracted from the pool5 layer across 20 16-frame clips and averaged) with linear SVM classifiers on UCF101 (3 splits) and ActivityNet (v1.3 validation set).

    Dataset Method Top-1 Top-3 / AUC
    UCF101 C3D + linear SVM 82.3% –
    UCF101 ResNet-152 + linear SVM 83.5% –
    UCF101 P3D ResNet + linear SVM 88.6% –
    UCF101 P3D ResNet + IDT 93.7% –
    UCF101 P3D ResNet (rec5c) + FV-VAE 90.5% –
    ActivityNet IDT 64.70% 77.98% (mAP: 68.69%)
    ActivityNet C3D 65.80% 81.16% (mAP: 67.68%)
    ActivityNet VGG 19 66.59% 82.70% (mAP: 70.22%)
    ActivityNet ResNet-152 71.43% 86.45% (mAP: 76.56%)
    ActivityNet P3D ResNet 75.12% 87.71% (mAP: 78.86%)

    On UCF101, single-modality RGB frame-based P3D ResNet achieves 88.6%, outperforming TSN using RGB frames (85.7%). On ActivityNet, P3D ResNet outperforms ResNet-152 by 3.7% and C3D by 9.3% in Top-1 accuracy.

  7. Knowl 7 — Action Similarity Labeling on the ASLAN Benchmark

    data/table

    Evaluation on the ASLAN dataset (3,697 videos across 432 action categories, 10-fold cross-validation) for pair-wise action similarity verification. For each 16-frame clip, representations are extracted from four layers of P3D ResNet: prob, pool5, res5c, and res4b35. For each layer, 12 similarity metrics are computed per video pair, resulting in a 48-dimensional similarity feature classified via linear SVM after L2L_2 normalization.

    Method Model Accuracy AUC
    STIP linear 60.9% 65.3%
    MIP metric 65.5% 71.9%
    IDT + FV metric 68.7% 75.4%
    C3D linear 78.3% 86.5%
    ResNet-152 linear 70.4% 77.4%
    P3D ResNet linear 80.8% 87.9%

    P3D ResNet outperforms C3D by 2.5% in accuracy and 1.4% in AUC, and outperforms ResNet-152 by 10.4% in accuracy.

  8. Knowl 8 — Video Scene Recognition Performance on Dynamic Scene and YUPENN

    data/table

    Evaluation of spatio-temporal representations for scene recognition on Dynamic Scene (13 categories, 10 videos per category) and YUPENN (14 categories, 30 videos per category) using the standard leave-one-video-out protocol.

    Method Dynamic Scene Accuracy YUPENN Accuracy
    Derpanis et al. 43.1% 80.7%
    Theriault et al. 74.6% 85.0%
    Feichtenhofer et al. 77.7% 96.2%
    C3D 87.7% 98.1%
    ResNet-152 93.6% 99.2%
    P3D ResNet 94.6% 99.5%

    P3D ResNet achieves 94.6% on Dynamic Scene and 99.5% on YUPENN, outperforming both 3D CNN (C3D) and 2D deep features (ResNet-152).

  9. Knowl 9 — Robustness of P3D Representations to Dimensionality Reduction

    empirical result

    When original video feature vectors are compressed via Principal Component Analysis (PCA) to lower dimensions (d∈{500,300,200,100,50,10}d \in \{500, 300, 200, 100, 50, 10\}), P3D ResNet features consistently achieve higher action recognition accuracy on UCF101 than IDT, C3D, and ResNet-152 across all dimensions.

    While ResNet-152 accuracy drops steeply as dimensionality decreases (due to the domain gap of being pre-trained solely on 2D images without learning temporal representations from video), P3D ResNet features demonstrate stability and graceful degradation similar to 3D CNNs (C3D), indicating that combining pre-trained 2D spatial filters with learned 1D temporal connections produces compact, robust representations.

Coverage note — None omitted; all core architectural designs, building block formulations, training protocols, and benchmark results across action recognition, similarity labeling, and scene recognition are represented.

References

  1. 1.Deep draw. https://github.com/auduno/deepdraw. 5, 6
  2. 2.F. Caba Heilbron, V. Escorcia, B. Ghanem, and J. Carlos Niebles. Activitynet: A large-scale video benchmark for human activity understanding. In CVPR, 2015. 6
  3. 3.K. Derpanis, M. Lecce, K. Daniildis, and R. Wildes. Dynamic scene understanding: The role of orientation features in space and time in scene classification. In CVPR, 2012. 6, 8
  4. 4.C. Feichtenhofer, A. Pinz, and R. Wildes. Spatiotemporal residual networks for video action recognition. In NIPS, 2016. 7
  5. 5.C. Feichtenhofer, A. Pinz, and R. P. Wildes. Bags of spacetime energies for dynamic scene recognition. In CVPR, 2014. 8
  6. 6.C. Feichtenhofer, A. Pinz, and A. Zisserman. Convolutional two-stream network fusion for video action recognition. In CVPR, 2016. 2, 7
  7. 7.K. He, X. Zhang, S. Ren, and J. Sun. Deep residual learning for image recognition. In CVPR, 2016. 1, 3, 4, 5, 6, 7, 8
  8. 8.S. Ji, W. Xu, M. Yang, and K. Yu. 3d convolutional neural networks for human action recognition. IEEE Trans. on PAMI, 35, 2013. 1, 2, 3
  9. 9.Y.-G. Jiang, Z. Wu, J. Wang, X. Xue, and S.-F. Chang. Exploiting feature and class relationships in video categorization with regularized deep neural networks. IEEE Trans. on PAMI, 2017. 2
  10. 10.A. Karpathy, G. Toderici, S. Shetty, T. Leung, R. Sukthankar, and L. Fei-Fei. Large-scale video classification with convolutional neural networks. In CVPR, 2014. 1, 2, 5, 6
  11. 11.A. Klaser, M. Marszałek, and C. Schmid. A spatio-temporal descriptor based on 3d-gradients. In BMVC, 2008. 2
  12. 12.O. Kliper-Gross, Y. Gurovich, T. Hassner, and L. Wolf. Motion interchange patterns for action recognition in unconstrained videos. In ECCV, 2012. 7
  13. 13.O. Kliper-Gross, T. Hassner, and L. Wolf. The action similarity labeling challenge. IEEE Trans. on PAMI, 34(3), 2012. 6, 7
  14. 14.A. Krizhevsky, I. Sutskever, and G. E. Hinton. Imagenet classification with deep convolutional neural networks. In NIPS, 2012. 5
  15. 15.I. Laptev. On space-time interest points. International journal of computer vision, 64(2-3):107–123, 2005. 2
  16. 16.I. Laptev, M. Marszalek, C. Schmid, and B. Rozenfeld. Learning realistic human actions from movies. In CVPR, 2008. 2
  17. 17.Q. Li, Z. Qiu, T. Yao, T. Mei, Y. Rui, and J. Luo. Action recognition by learning deep multi-granular spatio-temporal video representation. In ICMR, 2016. 2
  18. 18.Q. Li, Z. Qiu, T. Yao, T. Mei, Y. Rui, and J. Luo. Learning hierarchical video representation for action recognition. International Journal of Multimedia Information Retrieval, pages 1–14, 2017. 2
  19. 19.X. Peng, Y. Qiao, Q. Peng, and Q. Wang. Large margin dimensionality reduction for action similarity labeling. IEEE Signal Processing Letters, 21(8):1022–1025, 2014. 7, 8
  20. 20.F. Perronnin, J. Sanchez, and T. Mensink. Improving the ´ fisher kernel for large-scale image classification. In ECCV, 2010. 2
  21. 21.Z. Qiu, Q. Li, T. Yao, T. Mei, and Y. Rui. Msr asia msm at thumos challenge 2015. In CVPR workshop, 2015. 2
  22. 22.Z. Qiu, T. Yao, and T. Mei. Deep quantization: Encoding convolutional activations with deep generative model. In CVPR, 2017. 7
  23. 23.P. Scovanner, S. Ali, and M. Shah. A 3-dimensional sift descriptor and its application to action recognition. In ACM MM, 2007. 2
  24. 24.N. Shroff, P. Turaga, and R. Chellappa. Moving vistas: Exploiting motion for describing scenes. In CVPR, 2010. 6
  25. 25.K. Simonyan and A. Zisserman. Two-stream convolutional networks for action recognition in videos. In NIPS, 2014. 2, 7
  26. 26.K. Simonyan and A. Zisserman. Very deep convolutional networks for large-scale image recognition. In ICLR, 2015. 7
  27. 27.K. Soomro, A. R. Zamir, and M. Shah. UCF101: A dataset of 101 human action classes from videos in the wild. CRCVTR-12-01, 2012. 4, 6
  28. 28.N. Srivastava, E. Mansimov, and R. Salakhutdinov. Unsupervised learning of video representations using lstms. In ICML, 2015. 2
  29. 29.L. Sun, K. Jia, D.-Y. Yeung, and B. E. Shi. Human action recognition using factorized spatio-temporal convolutional networks. In ICCV, 2015. 7
  30. 30.C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich. Going deeper with convolutions. In CVPR, 2015. 5
  31. 31.C. Theriault, N. Thome, and M. Cord. Dynamic scene classification: Learning motion descriptors with slow features analysis. In CVPR, 2013. 8
  32. 32.D. Tran, L. Bourdev, R. Fergus, L. Torresani, and M. Paluri. Learning spatiotemporal features with 3d convolutional networks. In ICCV, 2015. 1, 2, 3, 5, 6, 7, 8
  33. 33.L. van der Maaten and G. Hinton. Visualizing data using t-sne. JMLR, 2008. 8
  34. 34.G. Varol, I. Laptev, and C. Schmid. Long-term temporal convolutions for action recognition. In arXiv preprint arXiv:1604.04494, 2016. 1, 2, 7
  35. 35.H. Wang, A. Klaser, C. Schmid, and C.-L. Liu. Action recognition by dense trajectories. In CVPR, 2011. 2
  36. 36.H. Wang and C. Schmid. Action recognition with improved trajectories. In ICCV, 2013. 2, 7
  37. 37.L. Wang, Y. Qiao, and X. Tang. Action recognition with trajectory-pooled deep-convolutional descriptors. In CVPR, 2015. 2, 7
  38. 38.L. Wang, Y. Xiong, Z. Wang, Y. Qiao, D. Lin, X. Tang, and L. Van Gool. Temporal segment networks: towards good practices for deep action recognition. In ECCV, 2016. 2, 4, 5, 7
  39. 39.G. Willems, T. Tuytelaars, and L. Van Gool. An efficient dense and scale-invariant spatio-temporal interest point detector. In ECCV, 2008. 2
  40. 40.J. Yue-Hei Ng, M. Hausknecht, S. Vijayanarasimhan, O. Vinyals, R. Monga, and G. Toderici. Beyond short snippets: Deep networks for video classification. In CVPR, 2015. 2, 5, 6, 7
  41. 41.X. Zhang, Z. Li, C. C. Loy, and D. Lin. Polynet: A pursuit of structural diversity in very deep networks. arXiv preprint arXiv:1611.05725, 2016. 5
  42. 42.W. Zhu, J. Hu, G. Sun, X. Cao, and Y. Qiao. A key volume mining deep framework for action recognition. In CVPR, 2016. 2, 7

Citation

MLA
Qiu, Z., et al. “Learning Spatio-Temporal Representation with Pseudo-3D Residual Networks”. arXiv, 2017, http://arxiv.org/abs/1711.10305v1.
APA
Qiu, Z., Yao, T., & Mei, T. (2017). Learning Spatio-Temporal Representation with Pseudo-3D Residual Networks. arXiv. http://arxiv.org/abs/1711.10305v1
Chicago
Qiu, Z., T. Yao, and T. Mei. 2017. “Learning Spatio-Temporal Representation with Pseudo-3D Residual Networks”. arXiv. http://arxiv.org/abs/1711.10305v1.
Harvard
Qiu, Z., Yao, T. and Mei, T. (2017) “Learning Spatio-Temporal Representation with Pseudo-3D Residual Networks”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1711.10305v1.
Vancouver
1. Qiu Z, Yao T, Mei T (2017) Learning Spatio-Temporal Representation with Pseudo-3D Residual Networks. arXiv

BibTeX

@article{qiu2017learning,
  title = {Learning Spatio-Temporal Representation with Pseudo-3D Residual Networks},
  author = {Qiu, Zhaofan and Yao, Ting and Mei, Tao},
  year = {2017},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1711.10305v1},
  eprint = {1711.10305}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE