Abnormal Event Detection at 150 FPS in MATLAB
Cewu LuJianping ShiJiaya Jia
Proposes a sparse combination learning framework that replaces costly sparse coding with small-scale least-squares projections, enabling abnormal event detection in surveillance video at 150 frames per second in MATLAB without sacrificing accuracy.
Surveillance systems generate vast amounts of video data, yet manual monitoring is labor-intensive and inefficient because abnormal events occur rarely. While automated detection using sparse representation achieves high accuracy, existing methods are computationally intensive and require seconds to process a single frame. This computational bottleneck causes significant response delays and prevents real-time deployment on standard hardware.
The article evaluates a sparse combination learning framework designed to detect abnormal events in surveillance video at real-time speeds while preserving high detection accuracy.
The proposed method replaces complex per-frame optimization with a set of pre-learned sparse basis combinations. Video frames are decomposed into multi-scale spatial-temporal cubes, from which three-dimensional gradient features are extracted and compressed. During training, the system iteratively learns compact basis combinations bounded by a maximum reconstruction error. During testing, the system assesses incoming video features via simple matrix projections to identify anomalies. The authors validated the method on over 107 hours of surveillance video (spanning 31,200 feature groups) and benchmarked it on three public datasets: Avenue, Subway (Exit and Entrance gates), and UCSD Ped1.
The evaluation yielded several key findings. First, the framework achieved processing speeds of 140 to 150 frames per second on standard desktop hardware using MATLAB, representing a speed improvement of over 400 times compared to prior sparsity-based techniques (which took 2 to 4.6 seconds per frame). Second, surveillance video exhibited high structural redundancy; approximately 10 basis combinations per region were sufficient to represent normal activity, with 99% of regions requiring fewer than 45 combinations. Third, detection accuracy remained highly competitive. On the UCSD Ped1 dataset, the model achieved a 15% frame-level equal error rate (outperforming alternative sparse coding and subspace clustering baselines) and an area under the ROC curve of 91.8%. On the Subway and Avenue datasets, it maintained high detection rates with low false alarm counts (e.g., detecting 19 out of 19 ground-truth events at the Subway Exit Gate with only 2 false alarms).
These results demonstrate that automated surveillance can achieve true real-time processing without requiring specialized supercomputing hardware or sacrificing accuracy. By drastically cutting per-frame computation, security systems can issue immediate alerts and scale across multiple camera feeds on standard infrastructure, lowering hardware and operational costs.
Organizations seeking to implement or upgrade automated surveillance should consider adopting sparse combination structures over traditional dictionary-searching methods. Future engineering efforts should focus on extending this framework to other video analysis domains and implementing parallel processing pipelines to further minimize latency.
Confidence in these findings is supported by extensive validation across multiple diverse benchmark datasets and long-duration video feeds. However, stakeholders should note that the training phase relies on a sufficient volume of normal activity to establish reliable baseline representations. Unusual normal events not captured in the training footage may initially trigger false alarms until the baseline combination sets are updated.
- Paper: Efficient sparse coding algorithms, Honglak Lee et al. (2006). Introduces efficient convex optimization and feature-sign search methods for sparse coding that provide foundational algorithmic principles for fast sparse dictionary reconstruction.
- Paper: Online dictionary learning for sparse coding, Julien Mairal et al. (2009). Develops the online dictionary learning paradigm for sparse coding upon which real-time sparse combination frameworks for video modeling build.
- Paper: Robust Face Recognition via Sparse Representation, John Wright et al. (2009). Establishes sparse representation and reconstruction error as a robust formulation for pattern recognition and outlier detection.
- Paper: Learning Fast Approximations of Sparse Coding, Karol Gregor et al. (2010). Pioneers learned approximations of sparse coding to bypass iterative optimization bottlenecks and achieve real-time inference throughput.
- Paper: Robust Principal Component Analysis: Exact Recovery of Corrupted Low-Rank Matrices via Convex Optimization, John Wright et al. (2009). Provides the foundational mathematical framework for separating low-rank regular video structures from sparse anomalous variations.
- Paper: Action recognition by dense trajectories, Heng Wang et al. (2011). Defines standard spatio-temporal motion representations and feature descriptors widely utilized in video surveillance event analysis.
- Paper: Learning Temporal Regularity in Video Sequences, Mahmudul Hasan et al. (2016). Extends video anomaly detection by replacing handcrafted sparse dictionary models with deep convolutional autoencoders that learn temporal regularity end-to-end.
- Paper: Future Frame Prediction for Anomaly Detection - A New Baseline, Wen Liu et al. (2017). Advances surveillance anomaly detection beyond reconstruction-based formulations to predictive future-frame modeling using deep neural networks.
- Paper: Real-World Anomaly Detection in Surveillance Videos, Waqas Sultani et al. (2018). Generalizes real-time surveillance anomaly detection to weakly supervised deep multiple instance learning across large-scale untrimmed video benchmarks.
