Detecting Spacecraft Anomalies Using LSTMs and Nonparametric Dynamic Thresholding
Kyle HundmanValentino ConstantinouChristopher LaporteIan ColwellTom Soderstrom
Demonstrates an automated spacecraft anomaly detection framework that pairs Long Short-Term Memory networks with nonparametric dynamic thresholding to reliably flag telemetry issues and mitigate false alarms using real NASA mission data.
Modern spacecraft generate massive streams of performance data, creating an operational challenge for human engineers who must monitor thousands of telemetry channels for unexpected behavior. Traditional monitoring relies heavily on fixed, out-of-limits thresholds and manual chart reviews. These legacy systems require costly expert maintenance, struggle to scale with growing data volumes, and routinely fail to detect subtle anomalies that depend on time or operating context. To address these operational risks and reduce monitoring workloads, the article evaluated an automated anomaly detection framework designed to handle large-scale, complex telemetry data.
The evaluated approach uses Long Short-Term Memory neural networks, a form of recurrent neural network specialized for sequence data, to predict normal channel values one step ahead based on past readings and spacecraft command inputs. Prediction discrepancies are smoothed, and the framework applies a newly developed unsupervised, non-parametric dynamic thresholding technique to identify anomalous deviations without making flawed assumptions about standard data distributions. An anomaly pruning procedure was also introduced to filter out minor, noisy spikes and suppress false alarms. The authors validated this pipeline using historical, expert-labeled incident reports across 82 telemetry channels from two distinct missions: the Soil Moisture Active Passive satellite and the Mars Science Laboratory Curiosity rover.
The findings demonstrate that this combined framework significantly outperforms conventional statistical thresholding methods. The proposed dynamic thresholding with pruning achieved an overall precision of 87.5% and a recall of 80.0%, yielding the highest overall accuracy score across all tested methods. Pruning proved critical, boosting precision by nearly 39 percentage points with only a minimal 4.8 percentage point drop in recall. In contrast, standard Gaussian thresholding struggled because real-world prediction errors violated normal distribution assumptions. Furthermore, the approach successfully identified complex contextual anomalies—which represented 41% of all evaluated anomalies and are typically missed by traditional limit checks—achieving a 69.0% recall on contextual events and a 90.3% recall on point anomalies.
These results show that neural network forecasting paired with dynamic thresholding can reliably capture complex temporal failures while maintaining channel-level interpretability for engineering teams. However, performance varied by mission type; the routine operations of the satellite yielded higher accuracy (85.5% precision and recall) than the highly diverse, irregular activity sequences of the Mars rover (92.6% precision and 69.4% recall). An initial pilot deployment monitoring over 700 satellite channels confirmed several real-world anomalies, but it also underscored that suppressing false positives is essential to gain operational trust when screening hundreds of thousands of daily data points.
To move toward full operational adoption, mission teams should refine feature engineering by integrating detailed command context and event logs rather than relying solely on high-level command indicators. Operations should also incorporate human feedback loops to establish baseline alert scores for noisy channels and explore automated correlation tools to track inter-channel dependencies. While the experimental setup evaluated focused five-day spans around known incidents rather than multi-year continuous operations, the evidence provides strong confidence that this unsupervised framework offers a viable, scalable alternative to manual limit-setting in mission-critical environments.
- Paper: Long Short-Term Memory, Sepp Hochreiter et al. (1997). Introduces the Long Short-Term Memory architecture that forms the core sequence modeling and telemetry prediction engine used by the source paper.
- Paper: A Critical Review of Recurrent Neural Networks for Sequence Learning, Zachary C. Lipton et al. (2015). Provides a comprehensive foundational review of recurrent architectures, gated memory units, and sequence-learning dynamics essential for understanding recurrent time-series modeling.
- Paper: Isolation-Based Anomaly Detection, Fei Tony Liu et al. (2012). Establishes standard principles and baselines of unsupervised anomaly detection in complex, high-dimensional operational datasets.
- Paper: Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling, Junyoung Chung et al. (2014). Compares gated recurrent architectures to guide design choices in modeling temporal dependencies within sequential engineering data.
- Paper: Recurrent Neural Network Regularization, Wojciech Zaremba et al. (2014). Presents regularization and dropout strategies for recurrent neural networks that prevent overfitting during sequence training.
- Paper: Model Assertions for Monitoring and Improving ML Models, Daniel Kang et al. (2020). Extends continuous telemetry monitoring by using domain-specific assertions and runtime quality assurance to detect high-confidence operational errors.
- Paper: Failing Loudly: An Empirical Study of Methods for Detecting Dataset Shift, Stephan Rabanser et al. (2019). Explores how to identify and characterize distribution and dataset shifts in production systems to mitigate silent failures.
- Paper: Data Validation for Machine Learning, Eric Breck et al. (2019). Generalizes runtime telemetry checks into full production data validation pipelines that monitor schema drift and batch anomalies.
- Paper: Deep Anomaly Detection with Outlier Exposure, Dan Hendrycks et al. (2019). Advances deep anomaly detection by leveraging outlier exposure during training to improve out-of-distribution detection reliability.
- Paper: Energy-based Out-of-distribution Detection, Weitang Liu et al. (2020). Proposes energy-based scoring frameworks for detecting out-of-distribution inputs as an alternative to thresholding predictive errors.
- Paper: Deep learning for time series classification: a review, Hassan Ismail Fawaz et al. (2018). Systematically reviews modern deep neural network architectures for time series classification and sequence representation.
