Unsupervised Learning of Human Action Categories Using Spatial-Temporal Words

Juan Carlos NieblesHongchen WangLi Fei-Fei

article2008IJCV1,938 citations

Proposes an unsupervised framework using space-time interest points and probabilistic Latent Semantic Analysis to automatically recognize and spatio-temporally localize multiple human actions in complex video sequences without manual annotations.

Listen

Automated recognition and localization of human activities in video footage are increasingly critical for applications such as video surveillance, automated summarization, and digital content indexing. However, computer vision systems routinely struggle with real-world complexities, including cluttered backgrounds, camera movement, and video sequences containing multiple simultaneous activities. Traditional methods typically rely on manual tracking or labor-intensive supervised training, which limits scalability and increases operational costs.

The article evaluates an unsupervised approach that automatically learns human action categories from unlabeled video sequences without human annotation. The authors demonstrate that this model can successfully categorize and spatially localize individual and multiple actions within complex, unconstrained video streams.

To achieve this, the approach extracts local space-time interest points from video frames using separable linear filters and groups them into a vocabulary of visual "video words." A probabilistic Latent Semantic Analysis (pLSA) model then automatically learns the probability distributions corresponding to latent human action categories. The framework was evaluated on the benchmark KTH human motion dataset (598 sequences spanning six action types across 25 subjects) and a figure skating dataset (32 sequences spanning three actions across seven subjects), using leave-one-out cross-validation alongside additional complex multi-action videos.

The findings show strong classification performance and practical flexibility. On the KTH dataset, the unsupervised model achieved an 81.50% recognition accuracy, slightly outperforming state-of-the-art fully supervised models (81.17% and 71.72%) and significantly outperforming volumetric feature approaches (62.96%). On the figure skating dataset, the framework achieved an 80.67% average accuracy across complex movements. The model successfully identified and localized multiple distinct actions within single video streams and tracked sequential actions over extended timeframes. Most classification errors occurred between visually similar motions, such as running versus jogging or boxing versus hand clapping.

These results demonstrate that supervised manual labeling is not necessary to achieve high-accuracy video action recognition. By eliminating the requirement for manual annotations and body tracking, this unsupervised framework significantly lowers implementation costs and operational overhead for large-scale video analytics systems.

Organizations planning to deploy automated video categorization should consider unsupervised bag-of-words architectures as viable alternatives to supervised classifiers. Before wide-scale adoption in mission-critical environments, practitioners should conduct pilot evaluations on domain-specific footage. Further research should focus on testing these models against larger, more varied video datasets to address limitations related to limited sample sizes and the visual ambiguity among closely related actions.

  • Paper: Unsupervised Learning of Video Representations using LSTMs, Nitish Srivastava et al. (2015). It carries unsupervised video learning beyond hand-engineered space-time words by learning temporal representations from video sequences for downstream action recognition.
  • Paper: Generating Videos with Scene Dynamics, Carl Vondrick et al. (2016). It extends the goal of learning from unlabeled video by modeling scene dynamics generatively and using the learned representations for action recognition.
Cover for Unsupervised Learning of Human Action Categories Using Spatial-Temporal Words

Abstract

We present a novel unsupervised learning method for human action categories. A video sequence is represented as a collection of spatial-temporal words by extracting space-time interest points. The algorithm automatically learns the probability distributions of the spatial-temporal words and intermediate topics corresponding to human action categories. This is achieved by using a probabilistic Latent Semantic Analysis (pLSA) model. Given a novel video sequence, the model can categorize and localize the human action(s) contained in the video. We test our algorithm on two challenging datasets: the KTH human action dataset and a recent dataset of figure skating actions. Our results are on par or slightly better than the best reported results. In addition, our algorithm can recognize and localize multiple actions in long and complex video sequences containing multiple motions.

Table of Contents

  • 1 Introduction
  • 2 Our Approach
  • 2.1 Feature Representation from Space-Time Interest Points
  • 2.2 Learning the Action Models: Latent Topic Discovery
  • 2.3 Categorization and Localization of Actions in a Testing Video
  • 3 Experimental Results
  • 3.1 Recognition and Localization of Single Actions
  • 3.1.1 Human Action Recognition and Localization Using KTH data
  • 3.1.2 Recognition and Localization of Figure Skating Actions
  • 3.2 Recognition and Localization of Multiple Actions in a Long Video Sequence
  • 4 Conclusion
  • References

Knowls

  1. Knowl 1 — Probabilistic Latent Semantic Analysis (pLSA) for Unsupervised Human Action Modeling

    model/method

    In the unsupervised action learning framework, a video collection is modeled using probabilistic Latent Semantic Analysis (pLSA). A video corpus comprises NN video sequences denoted by djd_j for j∈{1,…,N}j \in \{1, \dots, N\} and a vocabulary of MM spatial-temporal visual words wiw_i for i∈{1,…,M}i \in \{1, \dots, M\}. The corpus is summarized in an M×NM \times N co-occurrence table Nˉ\bar{N}, where n(wi,dj)n(w_i, d_j) denotes the number of occurrences of visual word wiw_i in video sequence djd_j.

    Each occurrence of a visual word wiw_i in video djd_j is associated with an unobserved latent topic variable zk∈{z1,…,zK}z_k \in \{z_1, \dots, z_K\}, where each topic corresponds to an action category and KK is the total number of action classes. The joint probability distribution over videos and visual words is given by: P(dj,wi)=P(dj)P(wi∣dj)P(d_j, w_i) = P(d_j) P(w_i | d_j)

    Assuming conditional independence of video documents and visual words given the latent topic, the conditional probability of observing visual word wiw_i in video djd_j is obtained by marginalizing over all KK latent topics: P(wi∣dj)=∑k=1KP(zk∣dj)P(wi∣zk)P(w_i | d_j) = \sum_{k=1}^K P(z_k | d_j) P(w_i | z_k) where P(zk∣dj)P(z_k | d_j) represents the probability distribution over latent action topics for video djd_j (video-specific mixture weights), and P(wi∣zk)P(w_i | z_k) represents the probability distribution over visual words for action topic zkz_k (shared across all videos).

    The model parameters are learned without supervision by maximizing the corpus log-likelihood using the Expectation-Maximization (EM) algorithm: L=∏i=1M∏j=1NP(wi∣dj)n(wi,dj)\mathcal{L} = \prod_{i=1}^M \prod_{j=1}^N P(w_i | d_j)^{n(w_i, d_j)}

  2. Knowl 2 — Space-Time Interest Point Detection and Feature Extraction

    model/method

    Video sequences are represented as collections of spatial-temporal visual words by detecting space-time interest points using separable linear filters applied to the video volume I(x,y,t)I(x,y,t). The spatio-temporal response function R(x,y,t)R(x,y,t) is defined as: R(x,y,t)=(I∗g(x,y;σ)∗hev(t;τ,ω))2+(I∗g(x,y;σ)∗hod(t;τ,ω))2R(x,y,t) = \left(I * g(x,y;\sigma) * h_{\text{ev}}(t;\tau,\omega)\right)^2 + \left(I * g(x,y;\sigma) * h_{\text{od}}(t;\tau,\omega)\right)^2 where:

    • g(x,y;σ)g(x,y;\sigma) is a 2D spatial Gaussian smoothing kernel applied along the spatial dimensions (x,y)(x,y) with spatial scale parameter σ\sigma.
    • hev(t;τ,ω)=−cos⁡(2πtω)e−t2/τ2h_{\text{ev}}(t;\tau,\omega) = -\cos(2\pi t \omega) e^{-t^2/\tau^2} and hod(t;τ,ω)=−sin⁡(2πtω)e−t2/τ2h_{\text{od}}(t;\tau,\omega) = -\sin(2\pi t \omega) e^{-t^2/\tau^2} are a quadrature pair of 1D temporal Gabor filters with temporal scale τ\tau and temporal frequency parameter ω=4/τ\omega = 4/\tau.

    Space-time interest points are identified at the local maxima of the response function RR. Around each detected point, a local 3D cuboid of dimensions approximately 6σ×6σ×6τ6\sigma \times 6\sigma \times 6\tau is extracted. For each cuboid, spatial-temporal brightness gradients are computed and concatenated into a descriptor vector, which is subsequently projected to a lower-dimensional space using Principal Component Analysis (PCA). A codebook of MM visual codewords is constructed by clustering a sampled set of these PCA-projected gradient descriptors using kk-means.

  3. Knowl 3 — Inference and Action Categorization for Novel Videos

    model/method

    Given learned action-topic conditional word distributions P(wi∣zk)P(w_i | z_k) for visual words wi∈{w1,…,wM}w_i \in \{w_1, \dots, w_M\} and latent action topics zk∈{z1,…,zK}z_k \in \{z_1, \dots, z_K\}, an unseen testing video dtestd_{\text{test}} is categorized by inferring its mixture coefficients P(zk∣dtest)P(z_k | d_{\text{test}}).

    The novel video is represented by its empirical word frequency distribution P~(w∣dtest)\tilde{P}(w | d_{\text{test}}), where P~(wi∣dtest)=n(wi,dtest)∑m=1Mn(wm,dtest)\tilde{P}(w_i | d_{\text{test}}) = \frac{n(w_i, d_{\text{test}})}{\sum_{m=1}^M n(w_m, d_{\text{test}})}.

    The mixture weights P(zk∣dtest)P(z_k | d_{\text{test}}) are estimated by projecting the empirical distribution onto the probability simplex spanned by the pre-learned distributions P(w∣zk)P(w | z_k), which minimizes the Kullback-Leibler (KL) divergence between P~(w∣dtest)\tilde{P}(w | d_{\text{test}}) and P(w∣dtest)=∑k=1KP(zk∣dtest)P(w∣zk)P(w | d_{\text{test}}) = \sum_{k=1}^K P(z_k | d_{\text{test}}) P(w | z_k) via the Expectation-Maximization (EM) algorithm while keeping P(w∣zk)P(w | z_k) fixed.

    The testing video dtestd_{\text{test}} is assigned to the overall action category corresponding to the maximum mixture weight: k∗=arg⁡max⁡k∈{1,…,K}P(zk∣dtest)k^* = \arg\max_{k \in \{1, \dots, K\}} P(z_k | d_{\text{test}})

  4. Knowl 4 — Simultaneous Multi-Action Spatial Localization via Topic Posteriors

    algorithm

    To localize multiple actions within a single video sequence djd_j, the posterior probability of action category zkz_k given each visual word occurrence wiw_i is computed using Bayes' rule: P(zk∣wi,dj)=P(wi∣zk)P(zk∣dj)∑l=1KP(wi∣zl)P(zl∣dj)P(z_k | w_i, d_j) = \frac{P(w_i | z_k) P(z_k | d_j)}{\sum_{l=1}^K P(w_i | z_l) P(z_l | d_j)}

    Spatial localization and multi-action segmentation proceed as follows:

    Input: Video sequence djd_j, extracted interest points with coordinates and words (xm,ym,tm,wm)m=1Npts{(x_m, y_m, t_m, w_m)}_{m=1}^{N_{\text{pts}}}, learned distributions P(w∣z)P(w|z), video topic mixture P(z∣dj)P(z|d_j)
    Output: Action labels and bounding boxes for localized action instances
    1. For each interest point m=1,…,Nptsm = 1, \dots, N_{\text{pts}}:
        Compute topic posterior P(zk∣wm,dj)=P(wm∣zk)P(zk∣dj)∑l=1KP(wm∣zl)P(zl∣dj)P(z_k | w_m, d_j) = \frac{P(w_m | z_k) P(z_k | d_j)}{\sum_{l=1}^K P(w_m | z_l) P(z_l | d_j)} for each k∈{1,…,K}k \in \{1, \dots, K\}
        Assign point mm the label k∗(m)=arg⁡max⁡kP(zk∣wm,dj)k^*(m) = \arg\max_k P(z_k | w_m, d_j)
    2. Identify the number CC of significantly active action categories present in djd_j based on P(zk∣wm,dj)P(z_k | w_m, d_j)
    3. Run KK-means clustering with CC clusters on the 2D spatial positions (xm,ym)(x_m, y_m) of all interest points
    4. For each spatial cluster c∈{1,…,C}c \in \{1, \dots, C\}:
        Determine the action label of cluster cc by majority vote over the word labels k∗(m)k^*(m) in cluster cc
        Compute the principal axes and eigenvalues of the spatial coordinates in cluster cc
        Construct a bounding box aligned with the principal axes scaled by the eigenvalues
  5. Knowl 5 — Human Action Recognition Accuracy on the KTH Dataset

    empirical result

    The pLSA unsupervised action recognition method was evaluated on the KTH human action dataset (598 video sequences across 6 actions: walking, jogging, running, boxing, hand waving, hand clapping; 25 subjects in indoor and outdoor settings with scale variations). Codebook construction used ~60,000 randomly selected space-time patches from 2 videos per action across 3 subjects. Evaluation used leave-one-out cross-validation across the remaining 22 subjects (averaged over 25 runs, excluding codebook subjects from testing).

    Using 500 visual codewords, the unsupervised pLSA method achieved an average recognition accuracy of 81.50%, matching or exceeding supervised methods:

    Method Recognition Accuracy (%) Learning Supervision Multiple Actions
    pLSA (Our method) 81.50 Unlabeled Yes
    Dollár et al. 81.17 Labeled No
    Schuldt et al. 71.72 Labeled No
    Ke et al. 62.96 Labeled No

    The 6-class confusion matrix with 500 codewords was:

    • Walking: 79% correct (14% confused with jogging, 6% with hand waving, 1% with running)
    • Running: 88% correct (11% confused with jogging, 1% with walking)
    • Jogging: 52% correct (36% confused with running, 11% with walking, 1% with hand clapping)
    • Hand waving: 93% correct (6% confused with boxing, 1% with hand clapping)
    • Hand clapping: 77% correct (23% confused with boxing)
    • Boxing: 100% correct
  6. Knowl 6 — Action Recognition Performance on the Figure Skating Dataset

    empirical result

    The pLSA unsupervised learning approach was evaluated on a figure skating dataset consisting of 32 video sequences of 7 people performing 3 actions: stand-spin, sit-spin, and camel-spin, recorded under moving camera conditions, non-stationary background, and rapid motion. Codebook generation was performed on patches from 6 subjects, and performance was assessed using leave-one-out cross-validation across subjects over 7 runs.

    Using a codebook of 1200 visual codewords (chosen to prevent overfitting in the generative model), the method achieved an average recognition accuracy of 80.67%. The class-specific confusion matrix was:

    • Stand-spin: 83% correct (17% confused with camel-spin, 0% with sit-spin)
    • Sit-spin: 67% correct (33% confused with stand-spin, 0% with camel-spin)
    • Camel-spin: 92% correct (8% confused with sit-spin, 0% with stand-spin)
  7. Knowl 7 — Temporal Action Segmentation in Continuous Video Sequences via Sliding Windows

    model/method

    To identify and segment sequential human actions over long, continuous video sequences containing multiple distinct motions over time, a temporal sliding window is employed.

    For each frame tt in a long video sequence:

    1. A temporal sub-sequence window centered at frame tt is extracted.
    2. Space-time interest points within the window are mapped to codebook visual words to construct an empirical word distribution P~(w∣dwin(t))\tilde{P}(w | d_{\text{win}(t)}).
    3. The mixture distribution over action categories P(z∣dwin(t))P(z | d_{\text{win}(t)}) is inferred by projecting P~(w∣dwin(t))\tilde{P}(w | d_{\text{win}(t)}) onto the pre-learned category word distributions P(w∣z)P(w | z) using the Expectation-Maximization algorithm.
    4. Frame tt is assigned the action category label achieving the maximum posterior probability: z∗(t)=arg⁡max⁡k∈{1,…,K}P(zk∣dwin(t))z^*(t) = \arg\max_{k \in \{1, \dots, K\}} P(z_k | d_{\text{win}(t)})

Coverage note — Qualitative localization visualizations on the auxiliary Caltech dataset and the brief mention of Latent Dirichlet Allocation (LDA) performance being slightly lower than pLSA were omitted as minor supporting observations.

References

  1. 1.Ankur Agarwal and Bill Triggs. Learning to track 3d human motion from silhouettes. In International Conference on Machine Learning, pages 9–16, Banff, July 2004.
  2. 2.Moshe Blank, Lena Gorelick, Eli Shechtman, Michal Irani, and Ronen Basri. Actions as space-time shapes. In ICCV, pages 1395–1402, 2005.
  3. 3.David M. Blei, Andrew Y. Ng, and Michael I. Jordan. Latent dirichlet allocation. Journal of Machine Learning Research, (3):993–1022, 2003.
  4. 4.Oren Boiman and Michal Irani. Detecting irregularities in images and in video. In Proceedings of International Conference on Computer Vision, volume 1, pages 462–469, 2005.
  5. 5.Vincent Cheung, Brendan J. Frey, and Nebojsa Jojic. Video epitomes. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, volume 1, pages 42–49, 2005.
  6. 6.Chris Dance, Jutta Willamowski, Lixin Fan, Cedric Bray, and Gabriela Csurka. Visual categorization with bags of keypoints. In ECCV International Workshop on Statistical Learning in Computer Vision, 2004.
  7. 7.Piotr Dollár, Vincent Rabaud, Garrison Cottrell, and Serge Belongie. Behavior recognition via sparse spatio-temporal features. In VS-PETS 2005, pages 65–72, 2005.
  8. 8.Alexei A. Efros, Alexander C. Berg, Greg Mori, and Jitendra Malik. Recognizing action at a distance. In IEEE International Conference on Computer Vision, pages 726–733, Nice, France, 2003.
  9. 9.Claudio Fanti, Lihi Zelnik-Manor, and Pietro Perona. Hybrid models for human motion recognition. In ICCV, volume 1, pages 1166–1173, 2005.
  10. 10.Li Fei-Fei and Pietro Perona. A bayesian hierarchical model for learning natural scene categories. In Proceedings of Computer Vision and Pattern Recognition, pages 524–531, 2005.
  11. 11.Chris Harris and Mike Stephens. A combined corner and edge detector. In Alvey Vision Conferences, pages 147–152, 1988.
  12. 12.Thomas Hofmann. Probabilistic latent semantic indexing. In SIGIR, pages 50–57, August 1999.
  13. 13.Yan Ke, Rahul Sukthankar, and Martial Hebert. Efficient visual event detection using volumetric features. In International Conference on Computer Vision, pages 166–173, 2005.
  14. 14.Ivan Laptev and Tony Lindeberg. Space-time interest points. In Proceedings of the ninth IEEE International Conference on Computer Vision, volume 1, pages 432 – 439, 2003.
  15. 15.Deva Ramanan and David A. Forsyth. Automatic annotation of everyday movements. In Sebastian Thrun, Lawrence Saul, and Bernhard Schölkopf, editors, Advances in Neural Information Processing Systems 16. MIT Press, Cambridge, MA, 2004.
  16. 16.Cordelia Schmid, Roger Mohr, and Christian Bauckhage. Evaluation of interest point detectors. International Journal of Computer Vision, 2(37):151–172, 2000.
  17. 17.Christian Schuldt, Ivan Laptev, and Barbara Caputo. Recognizing human actions: A local svm approach. In ICPR, pages 32–36, 2004.
  18. 18.Josef Sivic, Bryan C. Russell, Alexei A. Efros, Andrew Zisserman, and William T. Freeman. Discovering objects and their location in images. In International Conference on Computer Vision (ICCV), pages 370 – 377, October 2005.
  19. 19.Yang Song, Luis Goncalves, and Pietro Perona. Unsupervised learning of human motion. IEEE Transactions on Pattern Analysis and Machine Intelligence, 25(25):1–14, 2003.
  20. 20.Yang Wang, Hao Jiang, Mark S. Drew, Ze-Nian Li, and Greg Mori. Unsupervised discovery of action classes. In CVPR, 2006.
  21. 21.Alper Yilmaz and Mubarak Shah. Recognizing human actions in videos acquired by uncalibrated moving cameras. In IEEE International Conf. on Computer Vision (ICCV), volume 1, pages 150 – 157, 2005.

Citation

MLA
Niebles, J. C., et al. “Unsupervised Learning of Human Action Categories Using Spatial-Temporal Words”. International Journal of Computer Vision, vol. 79, no. 3, 2008, pp. 299–318, https://doi.org/10.1007/s11263-007-0122-4.
APA
Niebles, J. C., Wang, H., & Fei-Fei, L. (2008). Unsupervised Learning of Human Action Categories Using Spatial-Temporal Words. International Journal of Computer Vision, 79(3), 299–318. https://doi.org/10.1007/s11263-007-0122-4
Chicago
Niebles, J. C., H. Wang, and L. Fei-Fei. 2008. “Unsupervised Learning of Human Action Categories Using Spatial-Temporal Words”. International Journal of Computer Vision 79 (3): 299–318. https://doi.org/10.1007/s11263-007-0122-4.
Harvard
Niebles, J.C., Wang, H. and Fei-Fei, L. (2008) “Unsupervised Learning of Human Action Categories Using Spatial-Temporal Words”, International Journal of Computer Vision, 79(3), pp. 299–318. Available at: https://doi.org/10.1007/s11263-007-0122-4.
Vancouver
1. Niebles JC, Wang H, Fei-Fei L (2008) Unsupervised Learning of Human Action Categories Using Spatial-Temporal Words. International Journal of Computer Vision 79:299–318

BibTeX

@article{Niebles_2008, title={Unsupervised Learning of Human Action Categories Using Spatial-Temporal Words}, volume={79}, ISSN={1573-1405}, url={http://dx.doi.org/10.1007/s11263-007-0122-4}, DOI={10.1007/s11263-007-0122-4}, number={3}, journal={International Journal of Computer Vision}, publisher={Springer Science and Business Media LLC}, author={Niebles, Juan Carlos and Wang, Hongcheng and Fei-Fei, Li}, year={2008}, month=Mar, pages={299–318} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF