Real-World Anomaly Detection in Surveillance Videos
Waqas SultaniChen ChenMubarak Shah
Presents a weakly-supervised deep multiple instance ranking approach and a large-scale dataset of 1,900 untrimmed surveillance videos to detect and temporally localize real-world anomalies without requiring labor-intensive clip-level annotations.
Public surveillance networks have expanded rapidly across urban centers, but human monitoring resources have not kept pace. Operators face an unsustainable ratio of cameras to human monitors, creating an urgent operational need for automated anomaly detection to flag critical incidents such as crimes and traffic accidents in real time. The article set out to evaluate a deep learning framework for detecting real-world anomalies in untrimmed surveillance videos using only video-level labels during training, avoiding the labor-intensive need for frame-by-frame annotations.
To develop and validate the framework, the authors built a benchmark dataset consisting of 1,900 untrimmed, real-world closed-circuit television videos totaling 128 hours across 13 realistic anomaly categories, such as fighting, robbery, and road accidents. The approach uses a weakly supervised multiple instance learning framework. Videos are divided into temporal segments, and a ranking loss function with sparsity and smoothness constraints trains a deep neural network to assign higher anomaly scores to anomalous segments than to normal ones, comparing the highest-scoring segments between positive and negative videos.
The findings show that the proposed method significantly outperforms existing approaches. First, the framework achieved an area under the receiver operating characteristic curve of 75.41%, substantially exceeding standard dictionary learning at 65.51%, deep autoencoders at 50.6%, and binary classifiers at 50.0%. Second, adding temporal smoothness and sparsity constraints improved the area under the curve from 74.44% to 75.41%, helping localize transient events accurately. Third, on normal surveillance footage, the framework reduced false alarm rates to 1.9%, compared to 3.1% for dictionary learning and 27.2% for deep autoencoders. Finally, testing existing action recognition models to classify specific anomalous activities on the new dataset yielded low baseline accuracies of 23.0% and 28.4%, underscoring the high complexity and real-world difficulty of untrimmed surveillance footage.
These results indicate that automated surveillance systems can achieve high detection rates and drastically reduce false alarms without requiring expensive, frame-level training data. Training on both normal and anomalous examples produces a more resilient model than relying solely on normal baseline models, which frequently misinterpret benign environmental changes as anomalies. Organizations adopting automated surveillance should consider weakly supervised ranking methods to cut annotation costs, but fine-grained activity classification requires further research before being deployed for automated categorization.
While the detection framework proves effective across diverse conditions, performance remains limited in challenging environments, such as dark scenes with low visibility, occlusions caused by insects, or sudden non-threatening crowd gatherings. Decision-makers can have high confidence in the core detection and localization performance under typical operating conditions, though operational deployments should maintain human-in-the-loop oversight to manage edge-case false alarms and illumination extremes.
- Paper: Two-Stream Convolutional Networks for Action Recognition in Videos, Karen Simonyan et al. (2014). This paper establishes the foundational two-stream convolutional architecture for video action recognition that the source model adapts to handle weakly labeled anomaly detection.
- Paper: Temporal Segment Networks: Towards Good Practices for Deep Action Recognition, Limin Wang et al. (2016). This work introduces temporal segment networks that efficiently model long-range video sequences, providing the temporal sampling foundation utilized by the source model.
- Paper: Isolation-Based Anomaly Detection, Fei Tony Liu et al. (2012). This paper presents isolation-based anomaly detection principles that inform the core objective of identifying rare irregularities in complex datasets.
- Paper: SlowFast Networks for Video Recognition, Christoph Feichtenhofer et al. (2018). This paper extends video recognition by introducing dual-pathway architectures that process spatial semantics and rapid motion, building directly upon the video understanding foundations established in the source.
- Paper: Deep Anomaly Detection with Outlier Exposure, Dan Hendrycks et al. (2019). This article advances anomaly detection capabilities through outlier exposure techniques, generalizing beyond the specific video anomaly ranking framework proposed in the source.
- Paper: Deep One-Class Classification, Lukas Ruff et al. (2018). This paper continues the investigation of anomaly detection by developing end-to-end deep one-class classification, following up on the broader challenge introduced by the source.
