Deep learning-based electroencephalography analysis: a systematic review
Yannick RoyHubert BanvilleIsabela AlbuquerqueAlexandre GramfortTiago H. FalkJocelyn Faubert
Synthesizes 156 studies applying deep learning to electroencephalography, quantifying a median 5.4% accuracy improvement over traditional baselines while exposing critical reproducibility deficits and providing practical guidelines for future research.
Electroencephalography (EEG), which records electrical brain activity via scalp sensors, is critical for diagnosing neurological conditions such as epilepsy and sleep disorders, as well as for operating brain–computer interfaces. However, traditional EEG processing relies heavily on manual artifact removal and specialized, hand-crafted feature extraction. These conventional methods are labor-intensive, struggle to generalize across different individuals due to high subject variability, and often fail to scale efficiently. Advanced machine learning techniques, particularly deep learning, offer a promising alternative by learning representations directly from data, but their true performance advantage and best practices have remained unclear.
The article provides a comprehensive evaluation of deep learning applied to non-invasive EEG analysis. It aims to determine whether deep learning architectures outperform traditional processing approaches, identify methodological trends across application domains, and establish standards for experimental reproducibility.
To conduct this evaluation, the authors performed a systematic review of 154 studies published between January 2010 and July 2018 across scientific journals, conference proceedings, and electronic preprint repositories. The analysis covered five primary application areas: sleep staging, seizure detection, brain–computer interfaces, cognitive and affective monitoring, and processing tool improvements. For each study, the review extracted and analyzed roughly 70 distinct parameters spanning data characteristics, preprocessing pipelines, model architectures, training procedures, comparative performance metrics, and reproducibility factors.
The review revealed several key findings across the literature. First, deep learning models demonstrated a modest but consistent performance advantage, achieving a median accuracy gain of 5.4% over traditional baseline methods across evaluated tasks. Second, convolutional neural networks emerged as the dominant architecture, used in 40% of studies, followed by recurrent neural networks and autoencoders at 13% each. Most successful models were relatively shallow, using between 3 and 10 layers, which contrasts with the much deeper networks common in computer vision. Third, end-to-end learning proved viable: 49% of studies successfully trained models on raw or minimally filtered EEG time series rather than hand-engineered frequency features, and 47% bypassed explicit artifact removal without sacrificing performance. Finally, severe reproducibility deficiencies were identified across the field. While 53% of studies used publicly available data, only 13% shared their source code, leaving just 8% of the reviewed studies fully reproducible.
These findings indicate that deep learning can reduce the cost, timeline, and domain expertise required to build EEG pipelines by automating feature extraction and artifact handling. However, the modest 5.4% median performance improvement suggests that deep learning is not yet a complete replacement for established methods. The lack of standardized benchmarks and common baselines—such as state-of-the-art Riemannian geometry classifiers—means reported gains may be inflated by comparisons against weak baselines or compromised by publication bias toward positive results. Furthermore, because deep models can inadvertently learn non-brain artifacts (like muscle or eye movements) when trained on raw data, deploying uninspected models in clinical or critical settings introduces significant operational and compliance risks.
Organizations and researchers pursuing deep learning for EEG should prioritize rigorous, reproducible development workflows. Practitioners should adopt standardized reporting checklists, evaluate models on public benchmark datasets, and consistently compare new architectures against strong, open-source baselines rather than simplistic models. Where subject data is limited, teams should implement data augmentation techniques (such as overlapping windows) and explore hybrid transfer learning—pre-training models across large multi-subject pools before fine-tuning them on specific users. Before making substantial deployment decisions in clinical or high-stakes environments, organizations should conduct targeted pilot studies with transparent model inspection techniques to ensure decisions are driven by genuine neural signals rather than noise.
Confidence in the overarching architectural trends is high, but comparative performance conclusions must be interpreted with caution. The analyzed literature exhibits considerable heterogeneity in dataset sizes (ranging from under 10 minutes to thousands of hours), subject numbers (a median of only 13 subjects), validation schemes, and baseline selections. Readers should account for these data gaps and the prevailing lack of code transparency when evaluating reported performance gains.
- Paper: EEGNet: a compact convolutional neural network for EEG-based brain–computer interfaces, Vernon J. Lawhern et al. (2016). Introduces EEGNet, a foundational compact convolutional architecture widely reviewed and evaluated in the systematic survey of deep learning for EEG.
- Paper: Independent Component Analysis of Electroencephalographic Data, Scott Makeig et al. (1995). Establishes independent component analysis for EEG signal processing, providing the classical signal-decomposition baseline essential for understanding deep learning preprocessing approaches.
- Paper: Deep learning for time series classification: a review, Hassan Ismail Fawaz et al. (2018). Provides a comprehensive benchmark and architectural comparison of deep learning models on time-series classification, which directly informs how neural networks process temporal EEG signals.
- Paper: Time series classification from scratch with deep neural networks: A strong baseline, Zhiguang Wang et al. (2016). Presents strong baseline deep convolutional and residual architectures for raw time-series classification, laying the groundwork for end-to-end EEG modeling.
- Paper: Long Short-Term Memory, Sepp Hochreiter et al. (1997). Introduces Long Short-Term Memory networks, which form the primary recurrent neural network architecture evaluated throughout the EEG review.
- Paper: Deep learning in neural networks: An overview, Juergen Schmidhuber (2014). Surveys the historical foundations and core architectures of deep learning, contextualizing the convolutional and recurrent models applied to EEG.
- Paper: Methods for interpreting and understanding deep neural networks, Grégoire Montavon et al. (2018). Details methods for interpreting deep neural networks, addressing the critical need for feature attribution and explainability highlighted in clinical and BCI EEG studies.
- Paper: ICLabel: An automated electroencephalographic independent component classifier, dataset, and website, Luca Pion-Tonachini et al. (2019). Applies deep artificial neural networks to automate the classification of EEG independent components on a large-scale crowdsourced dataset, directly advancing automated EEG preprocessing.
- Paper: 1D Convolutional Neural Networks and Applications: A Survey, Serkan Kiranyaz et al. (2019). Synthesizes the specific theory and engineering deployment of lightweight 1D convolutional neural networks for real-time biological time-series analysis.
- Paper: Transformers in Time Series: A Survey, Qingsong Wen et al. (2022). Surveys Transformer adaptations for sequential and time-series data, representing the next major architectural evolution beyond the CNN and RNN paradigms reviewed in the source.
- Paper: Are we really making much progress? A worrying analysis of recent neural recommendation approaches, Maurizio Ferrari Dacrema et al. (2019). Investigates reproducibility issues and fair benchmarking against simple baselines in applied deep learning, extending the reproducibility critique raised in the EEG systematic review.
