Learning Temporal Regularity in Video Sequences

Mahmudul HasanJonghyun ChoiJan NeumannAmit K. Roy-ChowdhuryLarry S. Davis

article2016CVPR1,406 citations

Develops a fully convolutional autoencoder framework that learns regular spatio-temporal patterns with minimal supervision to detect anomalies in complex video sequences.

Listen

Modern video systems record vast quantities of footage, creating an operational burden where human reviewers must spend hours searching through uninformative scenes to spot critical events. Automated detection of unusual or meaningful activity remains difficult because anomalies are unpredictable and diverse, making standard supervised detection methods impractical. The article addresses this challenge by shifting the focus from identifying rare anomalies to modeling normal, recurring motion patterns—termed temporal regularity—using minimal supervision.

The main objective of the article is to demonstrate that an autoencoder framework can learn temporal regularity from ordinary video sequences and reliably detect anomalies by measuring reconstruction errors. Specifically, it evaluates two distinct architectures: a standard autoencoder trained on handcrafted motion trajectories and a fully convolutional autoencoder that learns both visual features and motion patterns directly from raw video clips.

The authors conducted experiments across several standard benchmark datasets, including CUHK Avenue, UCSD Pedestrian (Ped1 and Ped2), and Subway (Entrance and Exit) scenes, totaling nearly two hours of video. Instead of relying on manual event labeling, the models were trained under the assumption that training videos contain only regular activity. Credibility is further established by evaluating cross-dataset generalization without scene-specific fine-tuning.

The analysis produced several key findings. First, the fully convolutional autoencoder outperformed the handcrafted feature model across all benchmarks, achieving high detection rates such as 45 out of 47 anomalies on CUHK Avenue and 61 out of 66 on Subway Entrance. Second, the convolutional architecture successfully retained spatial context, enabling pixel-precise localization of irregular objects, such as vehicles on pedestrian paths or thrown items, whereas handcrafted features provided only coarse patch-level localization. Third, the model demonstrated strong generalization capabilities; training across multiple datasets simultaneously preserved performance without degrading accuracy on individual scenes. Fourth, the generative model proved capable of predicting short-range past and future regular frames from a single static image.

These findings suggest significant operational benefits for organizations managing surveillance, monitoring, or automated indexing pipelines. Shifting to an end-to-end unsupervised approach reduces the human cost and timeline associated with manual data labeling while improving computational efficiency compared to sparse-coding techniques. The system provides a unified generative model that adapts across varied surveillance camera angles and environments.

Stakeholders deploying automated surveillance should consider integrating fully convolutional autoencoders for anomaly triaging and video summarization. However, because the system flags any statistical deviation from regular motion, it generates higher false alarm rates when harmless but unusual actions occur (such as running in a subway). Decision-makers should implement these models alongside secondary filtering or human-in-the-loop validation rather than as fully autonomous alarms. Future efforts should evaluate the model on more complex, unconstrained outdoor environments and explore adaptive thresholding to mitigate false alarms.

arXiv: 1604.04574
Cover for Learning Temporal Regularity in Video Sequences

Abstract

Perceiving meaningful activities in a long video sequence is a challenging problem due to ambiguous definition of 'meaningfulness' as well as clutters in the scene. We approach this problem by learning a generative model for regular motion patterns, termed as regularity, using multiple sources with very limited supervision. Specifically, we propose two methods that are built upon the autoencoders for their ability to work with little to no supervision. We first leverage the conventional handcrafted spatio-temporal local features and learn a fully connected autoencoder on them. Second, we build a fully convolutional feed-forward autoencoder to learn both the local features and the classifiers as an end-to-end learning framework. Our model can capture the regularities from multiple datasets. We evaluate our methods in both qualitative and quantitative ways - showing the learned regularity of videos in various aspects and demonstrating competitive performance on anomaly detection datasets as an application.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Approach
  • 3.1 Learning Motions on Handcrafted Features
  • 3.1.1 Model Architecture
  • 3.2 Learning Features and Motions
  • 3.2.1 Model Architecture
  • 3.3 Optimization and Initialization
  • 3.4 Regularity Score
  • 4 Experiments
  • 4.1 Datasets
  • 4.2 Learning a General Model Across Datasets
  • 4.3 Visualizing Temporal Regularity
  • 4.4 Predicting the Regular Past and the Future
  • 4.5 Anomalous Event Detection
  • 4.6 Filter Responses
  • 5 Conclusion
  • References
  • 6 Dataset Details
  • 7 Learned Temporal Regularity
  • 7.1 CUHK Avenue Dataset
  • 7.2 UCSD Ped1
  • 7.3 UCSD Ped2
  • 7.4 Subway Enter
  • 7.5 Subway Exit
  • 8 Object Detection in Irregular Motion
  • 8.1 CUHK Avenue Dataset
  • 8.2 UCSD Ped1
  • 8.3 UCSD Ped2
  • 8.4 Subway Enter
  • 8.5 Subway Exit
  • 9 Predicting Past and Future Regular Frames
  • 9.1 CUHK Avenue Dataset
  • 9.2 UCSD Ped1
  • 9.3 UCSD Ped2
  • 9.4 Subway Enter
  • 9.5 Subway Exit
  • 10 Anomalous Event Detection and Generalization Analysis on Multiple Datasets
  • 10.1 CUHK Avenue Dataset
  • 10.2 UCSD Ped1
  • 10.3 UCSD Ped2
  • 10.4 Subway Enter
  • 10.5 Subway Exit
  • 11 Filter Response Visualization
  • 11.1 CUHK Avenue Dataset
  • 11.2 UCSD Ped1
  • 11.3 UCSD Ped2
  • 11.4 Subway Enter
  • 11.5 Subway Exit
  • 12 Filter Weights Visualization

Citation

MLA
Hasan, M., et al. “Learning Temporal Regularity in Video Sequences”. arXiv, 2016, http://arxiv.org/abs/1604.04574v1.
APA
Hasan, M., Choi, J., Neumann, J., Roy-Chowdhury, A. K., & Davis, L. S. (2016). Learning Temporal Regularity in Video Sequences. arXiv. http://arxiv.org/abs/1604.04574v1
Chicago
Hasan, M., J. Choi, J. Neumann, A. K. Roy-Chowdhury, and L. S. Davis. 2016. “Learning Temporal Regularity in Video Sequences”. arXiv. http://arxiv.org/abs/1604.04574v1.
Harvard
Hasan, M. et al. (2016) “Learning Temporal Regularity in Video Sequences”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1604.04574v1.
Vancouver
1. Hasan M, Choi J, Neumann J, Roy-Chowdhury AK, Davis LS (2016) Learning Temporal Regularity in Video Sequences. arXiv

BibTeX

@article{hasan2016learning,
  title = {Learning Temporal Regularity in Video Sequences},
  author = {Hasan, Mahmudul and Choi, Jonghyun and Neumann, Jan and Roy-Chowdhury, Amit K. and Davis, Larry S.},
  year = {2016},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1604.04574v1},
  eprint = {1604.04574}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE