Real-World Anomaly Detection in Surveillance Videos

Waqas SultaniChen ChenMubarak Shah

article2018CVPR2,171 citations

Presents a weakly-supervised deep multiple instance ranking approach and a large-scale dataset of 1,900 untrimmed surveillance videos to detect and temporally localize real-world anomalies without requiring labor-intensive clip-level annotations.

Listen

Public surveillance networks have expanded rapidly across urban centers, but human monitoring resources have not kept pace. Operators face an unsustainable ratio of cameras to human monitors, creating an urgent operational need for automated anomaly detection to flag critical incidents such as crimes and traffic accidents in real time. The article set out to evaluate a deep learning framework for detecting real-world anomalies in untrimmed surveillance videos using only video-level labels during training, avoiding the labor-intensive need for frame-by-frame annotations.

To develop and validate the framework, the authors built a benchmark dataset consisting of 1,900 untrimmed, real-world closed-circuit television videos totaling 128 hours across 13 realistic anomaly categories, such as fighting, robbery, and road accidents. The approach uses a weakly supervised multiple instance learning framework. Videos are divided into temporal segments, and a ranking loss function with sparsity and smoothness constraints trains a deep neural network to assign higher anomaly scores to anomalous segments than to normal ones, comparing the highest-scoring segments between positive and negative videos.

The findings show that the proposed method significantly outperforms existing approaches. First, the framework achieved an area under the receiver operating characteristic curve of 75.41%, substantially exceeding standard dictionary learning at 65.51%, deep autoencoders at 50.6%, and binary classifiers at 50.0%. Second, adding temporal smoothness and sparsity constraints improved the area under the curve from 74.44% to 75.41%, helping localize transient events accurately. Third, on normal surveillance footage, the framework reduced false alarm rates to 1.9%, compared to 3.1% for dictionary learning and 27.2% for deep autoencoders. Finally, testing existing action recognition models to classify specific anomalous activities on the new dataset yielded low baseline accuracies of 23.0% and 28.4%, underscoring the high complexity and real-world difficulty of untrimmed surveillance footage.

These results indicate that automated surveillance systems can achieve high detection rates and drastically reduce false alarms without requiring expensive, frame-level training data. Training on both normal and anomalous examples produces a more resilient model than relying solely on normal baseline models, which frequently misinterpret benign environmental changes as anomalies. Organizations adopting automated surveillance should consider weakly supervised ranking methods to cut annotation costs, but fine-grained activity classification requires further research before being deployed for automated categorization.

While the detection framework proves effective across diverse conditions, performance remains limited in challenging environments, such as dark scenes with low visibility, occlusions caused by insects, or sudden non-threatening crowd gatherings. Decision-makers can have high confidence in the core detection and localization performance under typical operating conditions, though operational deployments should maintain human-in-the-loop oversight to manage edge-case false alarms and illumination extremes.

  • Paper: SlowFast Networks for Video Recognition, Christoph Feichtenhofer et al. (2018). This paper extends video recognition by introducing dual-pathway architectures that process spatial semantics and rapid motion, building directly upon the video understanding foundations established in the source.
  • Paper: Deep Anomaly Detection with Outlier Exposure, Dan Hendrycks et al. (2019). This article advances anomaly detection capabilities through outlier exposure techniques, generalizing beyond the specific video anomaly ranking framework proposed in the source.
  • Paper: Deep One-Class Classification, Lukas Ruff et al. (2018). This paper continues the investigation of anomaly detection by developing end-to-end deep one-class classification, following up on the broader challenge introduced by the source.
Cover for Real-World Anomaly Detection in Surveillance Videos

Abstract

Surveillance videos are able to capture a variety of realistic anomalies. In this paper, we propose to learn anomalies by exploiting both normal and anomalous videos. To avoid annotating the anomalous segments or clips in training videos, which is very time consuming, we propose to learn anomaly through the deep multiple instance ranking framework by leveraging weakly labeled training videos, i.e. the training labels (anomalous or normal) are at video-level instead of clip-level. In our approach, we consider normal and anomalous videos as bags and video segments as instances in multiple instance learning (MIL), and automatically learn a deep anomaly ranking model that predicts high anomaly scores for anomalous video segments. Furthermore, we introduce sparsity and temporal smoothness constraints in the ranking loss function to better localize anomaly during training. We also introduce a new large-scale first of its kind dataset of 128 hours of videos. It consists of 1900 long and untrimmed real-world surveillance videos, with 13 realistic anomalies such as fighting, road accident, burglary, robbery, etc. as well as normal activities. This dataset can be used for two tasks. First, general anomaly detection considering all anomalies in one group and all normal activities in another group. Second, for recognizing each of 13 anomalous activities. Our experimental results show that our MIL method for anomaly detection achieves significant improvement on anomaly detection performance as compared to the state-of-the-art approaches. We provide the results of several recent deep learning baselines on anomalous activity recognition. The low recognition performance of these baselines reveals that our dataset is very challenging and opens more opportunities for future work. The dataset is available at: this https URL

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Proposed Anomaly Detection Method
  • 3.1 Multiple Instance Learning
  • 3.2 Deep MIL Ranking Model
  • 4 Dataset
  • 4.1 Previous datasets
  • 4.2 Our dataset
  • 5 Experiments
  • 5.1 Implementation Details
  • 5.2 Comparison with the State-of-the-art
  • 5.3 Analysis of the Proposed Method
  • 5.4 Anomalous Activity Recognition Experiments
  • 6 Conclusions
  • 7 Acknowledgement
  • References

Knowls

  1. Knowl 1 — Weakly Supervised Multiple Instance Ranking Formulation for Video Anomaly Detection

    model/method

    Video anomaly detection is formulated as a regression problem within a Multiple Instance Learning (MIL) ranking framework using weakly labeled training data. Videos are represented as bags and fixed-length non-overlapping temporal segments within a video serve as instances in the bag.

    Let an anomalous video containing an anomaly at an unknown temporal location be represented as a positive bag Ba=(Va1,Va2,…,Vam)B_a = (V_a^1, V_a^2, \dots, V_a^m), where mm is the number of video segments and at least one segment contains an anomaly. Let a normal video containing no anomalous behavior be represented as a negative bag Bn=(Vn1,Vn2,…,Vnm)B_n = (V_n^1, V_n^2, \dots, V_n^m), where all segments are normal.

    Rather than requiring segment-level ground-truth labels, the model is trained by enforcing that the instance with the highest anomaly score in the positive bag scores higher than the instance with the highest anomaly score in the negative bag:

    max⁡i∈Baf(Vai)>max⁡i∈Bnf(Vni)\max_{i \in B_a} f(V_a^i) > \max_{i \in B_n} f(V_n^i)

    where f(V)∈[0,1]f(V) \in [0, 1] represents the anomaly score predicted by a deep neural network for a video segment VV. The maximum-scoring instance in the positive bag represents the most likely anomalous segment, while the maximum-scoring instance in the negative bag acts as a hard negative example to suppress false alarms.

  2. Knowl 2 — MIL Anomaly Ranking Loss with Temporal Smoothness and Sparsity Constraints

    equation

    The loss function for training the anomaly detection network over a pair of positive bag BaB_a and negative bag BnB_n, each partitioned into nn temporal segments, is defined as:

    l(Ba,Bn)=max⁡(0,1−max⁡i∈Baf(Vai)+max⁡i∈Bnf(Vni))+λ1∑i=1n−1(f(Vai)−f(Vai+1))2+λ2∑i=1nf(Vai)l(B_a, B_n) = \max\left(0, 1 - \max_{i \in B_a} f(V_a^i) + \max_{i \in B_n} f(V_n^i)\right) + \lambda_1 \sum_{i=1}^{n-1} \left(f(V_a^i) - f(V_a^{i+1})\right)^2 + \lambda_2 \sum_{i=1}^n f(V_a^i)

    where:

    • f(Vai)f(V_a^i) is the predicted anomaly score for the ii-th temporal segment of the anomalous video bag BaB_a.
    • f(Vni)f(V_n^i) is the predicted anomaly score for the ii-th temporal segment of the normal video bag BnB_n.
    • The first term is the hinge ranking loss that enforces a margin of at least 11 between the highest scoring segment of BaB_a and the highest scoring segment of BnB_n.
    • The second term enforces temporal smoothness by penalizing score discrepancies between temporally adjacent segments in the anomalous video.
    • The third term enforces temporal sparsity on the anomaly scores of BaB_a, reflecting that anomalous events typically span only a small fraction of untrimmed surveillance videos.
    • λ1\lambda_1 and λ2\lambda_2 are regularization hyperparameters, both set to 8×10−58 \times 10^{-5} for optimal performance.

    The complete training objective regularized with model weights WW is:

    L(W)=l(Ba,Bn)+∥W∥F\mathcal{L}(W) = l(B_a, B_n) + \|W\|_F

    where ∥W∥F\|W\|_F denotes the Frobenius norm of the network weights.

  3. Knowl 3 — UCF-Crime Large-Scale Surveillance Video Anomaly Detection Dataset

    definition

    The UCF-Crime dataset is a large-scale benchmark of 1,900 untrimmed real-world CCTV surveillance videos totaling 128 hours (average video length of 7,247 frames). It covers 13 distinct realistic anomalous activities with significant public safety relevance: Abuse, Arrest, Arson, Assault, Accident, Burglary, Explosion, Fighting, Robbery, Shooting, Stealing, Shoplifting, and Vandalism, along with normal surveillance activities.

    The dataset is divided into:

    • Training set: 800 normal videos and 810 anomalous videos (Abuse: 48, Arrest: 45, Arson: 41, Assault: 47, Burglary: 87, Explosion: 29, Fighting: 45, Road Accidents: 127, Robbery: 145, Shooting: 27, Shoplifting: 29, Stealing: 95, Vandalism: 45).
    • Testing set: 150 normal videos and 140 anomalous videos with frame-level ground truth annotations obtained by averaging multiple annotators' temporal boundary labels.
    Dataset # of videos Average frames Dataset length Example anomalies
    UCSD Ped1 70 201 5 min Bikers, small carts, walking across walkways
    UCSD Ped2 28 163 5 min Bikers, small carts, walking across walkways
    Subway Entrance 1 121,749 1.5 hours Wrong direction, No payment
    Subway Exit 1 64,901 1.5 hours Wrong direction, No payment
    Avenue 37 839 30 min Run, throw, new object
    UMN 5 1290 5 min Run
    BOSS 12 4052 27 min Harass, Disease, Panic
    UCF-Crime (Ours) 1900 7247 128 hours Abuse, arrest, arson, assault, accident, burglary, fighting, robbery, etc.
  4. Knowl 4 — Deep Feature Extraction and Anomaly Scoring Network Architecture

    model/method

    The video anomaly detection pipeline operates as follows:

    1. Preprocessing and Video Partitioning: Video frames are resized to 240×320240 \times 320 pixels at a fixed frame rate of 30 frames per second. Each video is divided into n=32n = 32 non-overlapping temporal segments.
    2. Feature Extraction: Visual features are extracted from the FC6 layer of a pre-trained 3D Convolutional Network (C3D) for every 16-frame clip and normalized using L2L_2-normalization. The segment-level feature vector (4096-D) is computed by averaging the 16-frame clip descriptors within each segment.
    3. Fully Connected Scoring Network: The 4096-D segment feature vector is passed to a 3-layer fully connected network:
      • Layer 1: 512 units with Rectified Linear Unit (ReLU) activation.
      • Layer 2: 32 units.
      • Layer 3: 1 unit with Sigmoid activation producing a scalar anomaly score in [0,1][0, 1].
      • Dropout regularization with a 60% rate is applied between the FC layers.
    4. Optimization: The network is optimized using the Adagrad optimizer with an initial learning rate of 0.0010.001. Training mini-batches consist of 30 randomly selected positive bags and 30 negative bags.
  5. Knowl 5 — Frame-Level Anomaly Detection Performance Comparison on UCF-Crime

    data/table

    Performance is evaluated using the frame-level Receiver Operating Characteristic (ROC) curve and Area Under the Curve (AUC) on the 290 test videos (140 anomalous, 150 normal) of the UCF-Crime dataset.

    Method AUC (%)
    Binary classifier (Linear SVM on C3D) 50.00
    Hasan et al. (2016) Fully Convolutional Auto-Encoder 50.60
    Lu et al. (2013) Sparse Dictionary Learning 65.51
    Proposed Deep MIL (without smoothness and sparsity constraints) 74.44
    Proposed Deep MIL (with smoothness and sparsity constraints) 75.41

    The deep MIL ranking approach outperforms unsupervised dictionary reconstruction and auto-encoder baselines by a large margin (+9.9%+9.9\% AUC over sparse coding). Adding temporal smoothness and sparsity constraints yields a further gain of 0.97%0.97\% AUC.

  6. Knowl 6 — False Alarm Rate Comparison on Normal Testing Surveillance Videos

    data/table

    False alarm rates evaluated on normal surveillance test videos at an anomaly detection threshold of 50% (0.50 score):

    Method False Alarm Rate (%)
    Hasan et al. (2016) Auto-Encoder 27.2
    Lu et al. (2013) Sparse Dictionary 3.1
    Proposed Deep MIL Ranking 1.9

    By leveraging both normal and anomalous videos during weakly supervised training, the deep MIL ranking model learns more comprehensive and robust representations of normal activity patterns, achieving a substantially lower false alarm rate (1.9%1.9\%) compared to reconstruction-based methods (27.2%27.2\% and 3.1%3.1\%).

  7. Knowl 7 — Anomalous Activity Recognition Baselines on Untrimmed Videos

    data/table

    Multi-class anomalous activity recognition is benchmarked across the 13 anomalous classes of the UCF-Crime dataset using 50 videos per event class under a 4-fold cross-validation protocol (75% train / 25% test per fold).

    Method Classification Accuracy (%)
    C3D (4096-D clip-averaged feature vector + Nearest Neighbor) 23.0
    Tube Convolutional Neural Network (TCNN with Tube-of-Interest pooling) 28.4

    The low recognition accuracy achieved by state-of-the-art action recognition models underscores the difficulty of the benchmark, caused by long untrimmed sequences, low surveillance resolution, camera viewpoint shifts, illumination changes, and large intra-class variations.

  8. Knowl 8 — Failure Modes of Surveillance Video Anomaly Detection

    limitation

    The deep MIL anomaly detection model exhibits specific failure modes in challenging real-world surveillance scenarios:

    1. Severe Low-Light / Night Conditions: Anomalous events occurring in poorly illuminated environments (e.g., a nighttime office burglary where an intruder enters via a window) fail to be detected due to degraded visual feature quality and low contrast.
    2. Small Foreground Occlusions and Artifacts: High-frequency, non-event visual disturbances (such as flying insects directly in front of the camera lens) cause spurious activations and false alarms.
    3. Atypical but Normal Group Behaviors: Sudden, unusual gatherings of people (such as a crowd suddenly forming on a street to spectate a relay race) are misclassified as anomalies because the network fails to distinguish rare normal crowd gatherings from genuine disturbances.

Coverage note — No substantial contributed material was omitted; all primary methodological components, loss formulations, dataset specifications, detection benchmarks, false alarm analyses, activity recognition baselines, and failure modes are included.

References

  1. 1.http://www.multitel.be/image/research-development/research-projects/boss.php.
  2. 2.Unusual crowd activity dataset of university of minnesota. In http://mha.cs.umn.edu/movies/crowdactivity-all.avi.
  3. 3.A. Adam, E. Rivlin, I. Shimshoni, and D. Reinitz. Robust real-time unusual event detection using multiple fixed-location monitors. TPAMI, 2008.
  4. 4.S. Andrews, I. Tsochantaridis, and T. Hofmann. Support vector machines for multiple-instance learning. In NIPS, pages 577–584, Cambridge, MA, USA, 2002. MIT Press.
  5. 5.B. Anti and B. Ommer. Video parsing for abnormality detection. In ICCV, 2011.
  6. 6.R. Arandjelović, P. Gronat, A. Torii, T. Pajdla, and J. Sivic. NetVLAD: CNN architecture for weakly supervised place recognition. In CVPR, 2016.
  7. 7.A. Basharat, A. Gritai, and M. Shah. Learning object motion patterns for anomaly detection and improved object detection. In CVPR, 2008.
  8. 8.C. Bergeron, J. Zaretzki, C. Breneman, and K. P. Bennett. Multiple instance ranking. In ICML, 2008.
  9. 9.V. Chandola, A. Banerjee, and V. Kumar. Anomaly detection: A survey. ACM Comput. Surv., 2009.
  10. 10.X. Cui, Q. Liu, M. Gao, and D. N. Metaxas. Abnormal detection using interaction energy potentials. In CVPR, 2011.
  11. 11.A. Datta, M. Shah, and N. Da Vitoria Lobo. Person-on-person violence detection in video data. In ICPR, 2002.
  12. 12.T. G. Dietterich, R. H. Lathrop, and T. Lozano-Pérez. Solving the multiple instance problem with axis-parallel rectangles. Artificial Intelligence, 89(1):31–71, 1997.
  13. 13.S. Ding, L. Lin, G. Wang, and H. Chao. Deep feature learning with relative distance comparison for person re-identification. Pattern Recognition, 48(10):2993–3003, 2015.
  14. 14.J. Duchi, E. Hazan, and Y. Singer. Adaptive subgradient methods for online learning and stochastic optimization. J. Mach. Learn. Res., 2011.
  15. 15.Y. Gao, H. Liu, X. Sun, C. Wang, and Y. Liu. Violence detection using oriented violent flows. Image and Vision Computing, 2016.
  16. 16.A. Gordo, J. Almazań, J. Revaud, and D. Larlus. Deep image retrieval: Learning global representations for image search. In ECCV, 2016.
  17. 17.M. Gygli, Y. Song, and L. Cao. Video2gif: Automatic generation of animated gifs from video. In CVPR, June 2016.
  18. 18.M. Hasan, J. Choi, J. Neumann, A. K. Roy-Chowdhury, and L. S. Davis. Learning temporal regularity in video sequences. In CVPR, June 2016.
  19. 19.G. E. Hinton. Rectified linear units improve restricted boltzmann machines vinod nair. In ICML, 2010.
  20. 20.T. Hospedales, S. Gong, and T. Xiang. A markov clustering topic model for mining behaviour in video. In ICCV, 2009.
  21. 21.R. Hou, C. Chen, and M. Shah. Tube convolutional neural network (t-cnn) for action detection in videos. In ICCV, 2017.
  22. 22.T. Joachims. Optimizing search engines using clickthrough data. In ACM SIGKDD, 2002.
  23. 23.S. Kamijo, Y. Matsushita, K. Ikeuchi, and M. Sakauchi. Traffic monitoring and accident detection at intersections. IEEE Transactions on Intelligent Transportation Systems, 1(2):108–118, 2000.
  24. 24.A. Karpathy, G. Toderici, S. Shetty, T. Leung, R. Sukthankar, and L. Fei-Fei. Large-scale video classification with convolutional neural networks. In CVPR, 2014.
  25. 25.J. Kooij, M. Liem, J. Krijnders, T. Andringa, and D. Gavrila. Multi-modal human aggression detection. Computer Vision and Image Understanding, 2016.
  26. 26.L. Kratz and K. Nishino. Anomaly detection in extremely crowded scenes using spatio-temporal motion pattern models. In CVPR, 2009.
  27. 27.W. Li, V. Mahadevan, and N. Vasconcelos. Anomaly detection and localization in crowded scenes. TPAMI, 2014.
  28. 28.C. Lu, J. Shi, and J. Jia. Abnormal event detection at 150 fps in matlab. In ICCV, 2013.
  29. 29.R. Mehran, A. Oyama, and M. Shah. Abnormal crowd behavior detection using social force model. In CVPR, 2009.
  30. 30.S. Mohammadi, A. Perina, H. Kiani, and M. Vittorio. Angry crowds: Detecting violent events in videos. In ECCV, 2016.
  31. 31.I. Saleemi, K. Shafique, and M. Shah. Probabilistic modeling of scene dynamics for applications in visual surveillance. TPAMI, 31(8):1472–1485, 2009.
  32. 32.A. Sankaranarayanan, S. Alavi and R. Chellappa. Triplet similarity embedding for face verification. arXiv preprint arXiv:1602.03418, 2016.
  33. 33.N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov. Dropout: A simple way to prevent neural networks from overfitting. J. Mach. Learn. Res., 2014.
  34. 34.W. Sultani and J. Y. Choi. Abnormal traffic detection using intelligent driver model. In ICPR, 2010.
  35. 35.Theano Development Team. Theano: A Python framework for fast computation of mathematical expressions. arXiv preprint arXiv:1605.02688, 2016.
  36. 36.D. Tran, L. Bourdev, R. Fergus, L. Torresani, and M. Paluri. Learning spatiotemporal features with 3d convolutional networks. In ICCV, 2015.
  37. 37.J. Wang, Y. Song, T. Leung, C. Rosenberg, J. Wang, J. Philbin, B. Chen, and Y. Wu. Learning fine-grained image similarity with deep ranking. In CVPR, 2014.
  38. 38.S. Wu, B. E. Moore, and M. Shah. Chaotic invariants of lagrangian particle trajectories for anomaly detection in crowded scenes. In CVPR, 2010.
  39. 39.D. Xu, E. Ricci, Y. Yan, J. Song, and N. Sebe. Learning deep representations of appearance and motion for anomalous event detection. In BMVC, 2015.
  40. 40.T. Yao, T. Mei, and Y. Rui. Highlight detection with pairwise deep ranking for first-person video summarization. In CVPR, June 2016.
  41. 41.B. Zhao, L. Fei-Fei, and E. P. Xing. Online detection of unusual events in videos via dynamic sparse coding. In CVPR, pages 3313–3320, 2011.
  42. 42.B. Zhao, L. Fei-Fei, and E. P. Xing. Online detection of unusual events in videos via dynamic sparse coding. In CVPR, 2011.
  43. 43.Y. Zhu, I. M. Nayak, and A. K. Roy-Chowdhury. Context-aware activity recognition and anomaly detection in video. In IEEE Journal of Selected Topics in Signal Processing, 2013.

Citation

MLA
Sultani, W., et al. “Real-world Anomaly Detection in Surveillance Videos”. arXiv, 2018, http://arxiv.org/abs/1801.04264v3.
APA
Sultani, W., Chen, C., & Shah, M. (2018). Real-world Anomaly Detection in Surveillance Videos. arXiv. http://arxiv.org/abs/1801.04264v3
Chicago
Sultani, W., C. Chen, and M. Shah. 2018. “Real-world Anomaly Detection in Surveillance Videos”. arXiv. http://arxiv.org/abs/1801.04264v3.
Harvard
Sultani, W., Chen, C. and Shah, M. (2018) “Real-world Anomaly Detection in Surveillance Videos”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1801.04264v3.
Vancouver
1. Sultani W, Chen C, Shah M (2018) Real-world Anomaly Detection in Surveillance Videos. arXiv

BibTeX

@article{sultani2018real,
  title = {Real-world Anomaly Detection in Surveillance Videos},
  author = {Sultani, Waqas and Chen, Chen and Shah, Mubarak},
  year = {2018},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1801.04264v3},
  eprint = {1801.04264}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF

License: IEEE