InceptionTime: Finding AlexNet for time series classification

Hassan Ismail FawazBenjamin LucasGermain ForestierCharlotte PelletierDaniel F. SchmidtJonathan WeberGeoffrey I. WebbLhassane IdoumgharPierre-Alain MullerFrançois Petitjean

article2019Data mining and knowledge discovery1,771 citations

Introduces InceptionTime, an ensemble of deep convolutional neural networks that matches the state-of-the-art accuracy of HIVE-COTE for time series classification while training orders of magnitude faster and scaling to millions of sequences.

Listen

Modern organizations across healthcare, remote sensing, and activity recognition are generating unprecedented volumes of time series data. Extracting value from this data requires accurate and automated time series classification, which assigns categories to temporal sequences. While traditional algorithms such as HIVE-COTE deliver high classification accuracy, their extreme computational complexity makes training on massive, real-world datasets practically impossible.

The article introduces and evaluates InceptionTime, an ensemble of deep convolutional neural networks designed to achieve top-tier classification accuracy while drastically reducing training time. The study demonstrates the model's accuracy, scalability, and architectural design principles across standardized benchmarks and synthetic datasets.

To evaluate performance, the authors tested InceptionTime on 85 datasets from the public UCR time series archive and a synthetic dataset designed to isolate parameters such as sequence length and class count. InceptionTime combines five deep neural network models with identical architectures but different random weight initializations, using multiple parallel filters of varying lengths within specialized Inception modules to capture both short and long patterns.

The findings show that InceptionTime achieves classification accuracy on par with the class-leading HIVE-COTE algorithm, winning or tying on 46 of the 85 UCR benchmark datasets with no statistically significant difference in overall error. Crucially, InceptionTime is orders of magnitude faster: on a dataset of 1,500 short series, it trained in 1 hour compared to over 8 days for HIVE-COTE, and it successfully trained on 8 million series in 13 hours—a scale completely inaccessible to traditional methods. It also significantly outperforms the previous best deep learning ensemble, ResNet, winning 54 out of 85 dataset comparisons. Furthermore, architectural tests revealed that incorporating long filter lengths directly expands the model's receptive field to capture long temporal patterns effectively, while ensembling five models provides an optimal trade-off between variance reduction and computational overhead.

These results demonstrate that enterprises no longer need to compromise between high accuracy and computational feasibility in time series analytics. InceptionTime eliminates severe operational bottlenecks by leveraging standard graphics processing unit hardware, making large-scale automated monitoring, diagnostic tools, and predictive maintenance viable and cost-effective.

Organizations handling large-scale time series data should consider adopting InceptionTime as a scalable alternative to traditional ensembles. Data science teams should use an ensemble size of five models to balance training runtime and accuracy, employ bottleneck layers to cut network parameters by roughly half without sacrificing accuracy, and apply transfer learning with fine-tuning when working with small or specialized datasets like spectrography.

While confidence in the model's benchmark performance and scalability is high, the authors note that deeper architectures or excessively long filters can overfit very small datasets due to the limited number of labeled training samples in existing benchmarks. Stakeholders should remain cautious when deploying the model on very small datasets without validating against overfitting or utilizing transfer learning.

  • Paper: TimesNet: Temporal 2D-Variation Modeling for General Time Series Analysis, Haixu Wu et al. (2023). TimesNet builds upon multi-scale Inception-style architectures by transforming 1D series into 2D temporal variation maps to handle both classification and forecasting tasks.
  • Paper: Ensemble deep learning: A review, M. A. Ganaie et al. (2021). This survey provides a broader theoretical and empirical perspective on deep neural network ensembles, analyzing the diversity and variance-reduction mechanisms that underpin InceptionTime's ensembling strategy.
Cover for InceptionTime: Finding AlexNet for time series classification

Abstract

This paper brings deep learning at the forefront of research into Time Series Classification (TSC). TSC is the area of machine learning tasked with the categorization (or labelling) of time series. The last few decades of work in this area have led to significant progress in the accuracy of classifiers, with the state of the art now represented by the HIVE-COTE algorithm. While extremely accurate, HIVE-COTE cannot be applied to many real-world datasets because of its high training time complexity in O(N2 * T4) for a dataset with N time series of length T. For example, it takes HIVE-COTE more than 8 days to learn from a small dataset with N = 1500 time series of short length T = 46. Meanwhile deep learning has received enormous attention because of its high accuracy and scalability. Recent approaches to deep learning for TSC have been scalable, but less accurate than HIVE-COTE. We introduce InceptionTime - an ensemble of deep Convolutional Neural Network (CNN) models, inspired by the Inception-v4 architecture. Our experiments show that InceptionTime is on par with HIVE-COTE in terms of accuracy while being much more scalable: not only can it learn from 1,500 time series in one hour but it can also learn from 8M time series in 13 hours, a quantity of data that is fully out of reach of HIVE-COTE.

Table of Contents

  • 1 Introduction
  • 2 Related work
  • 2.1 Time series classification
  • 2.1.1 Whole series
  • 2.1.2 Dictionary based
  • 2.1.3 Shapelets
  • 2.1.4 Transformation ensembles
  • 2.2 Deep learning for time series classification
  • 3 InceptionTime: an accurate and scalable time series classifier
  • 3.1 Inception Network: a novel architecture for TSC
  • 3.2 InceptionTime: a neural network ensemble for TSC
  • 3.3 Receptive field
  • 4 Experimental setup
  • 5 Experiments: InceptionTime
  • 6 Architectural Hyperparameter study
  • 6.1 Batch size
  • 6.2 Bottleneck and residual connections
  • 6.3 Depth
  • 6.4 Filter length
  • 6.5 Number of filters
  • 6.6 Sensitivity analysis
  • 7 Conclusion
  • References

Knowls

  1. Knowl 1 — Inception Network and InceptionTime Ensemble Architecture

    model/method

    InceptionTime is an ensemble classifier for Time Series Classification (TSC) comprising n=5n = 5 independently trained, identical Inception networks with different random weight initializations.

    Each individual Inception network processes a multivariate or univariate time series input X∈RT×MX \in \mathbb{R}^{T \times M} of length TT and MM channels through the following sequence:

    1. Feature Extraction Backbone: Six stacked Inception modules arranged into two consecutive residual blocks (three Inception modules per block). A residual shortcut with a linear transformation connects the input of each residual block to its output via element-wise addition, mitigating vanishing gradients.

    2. Global Pooling: A Global Average Pooling (GAP) layer that averages each feature map across the entire time dimension TT, producing a single feature vector invariant to temporal position.

    3. Classification Head: A fully connected layer with a softmax activation containing KK output neurons, corresponding to the KK distinct target classes.

    Network weights are initialized using Glorot uniform initialization and optimized end-to-end with the Adam optimization algorithm using a default batch size of 64.

  2. Knowl 2 — One-Dimensional Inception Module Architecture

    model/method

    The Inception module is the core building block of the Inception network, designed to extract multi-resolution temporal features from an input multivariate time series (or intermediate feature map) with length TT and MM input channels:

    1. Bottleneck Layer: A 1D convolutional layer with mm filters of kernel length 1 and stride 1 (default m=32m = 32). It projects the input from MM channels to m≪Mm \ll M channels, reducing parameter count and computational complexity by nearly half.

    2. Parallel Multi-Scale Convolutions: Three parallel 1D convolutional branches applied simultaneously to the output of the bottleneck layer. Each branch applies 32 filters of length l∈{10,20,40}l \in \{10, 20, 40\} (with stride 1 and padding chosen to preserve sequence length TT).

    3. Parallel MaxPooling Branch: A MaxPool1D operation with window length 3 and stride 1 applied directly to the module's input, followed by a 1D bottleneck convolution with 32 filters of length 1 to reduce channel dimension and provide perturbation invariance.

    4. Feature Concatenation: The outputs of the three convolution branches and the MaxPooling branch are concatenated along the channel dimension, yielding an output representation with 32×4=12832 \times 4 = 128 feature channels of length TT.

  3. Knowl 3 — InceptionTime Ensemble Probability Aggregation

    equation

    For an input time series xix_i and a classification problem with CC classes, the ensemble prediction y^i,c\hat{y}_{i,c} for each class c∈{1,…,C}c \in \{1, \dots, C\} is computed as the arithmetic mean of the softmax outputs across nn randomly initialized Inception network models:

    y^i,c=1n∑j=1nσc(xi,θj)∀c∈[1,C]\hat{y}_{i,c} = \frac{1}{n} \sum_{j=1}^{n} \sigma_c(x_i, \theta_j) \quad \forall c \in [1, C]

    where θj\theta_j denotes the trained parameter weights of the jj-th Inception network instance, σc(xi,θj)\sigma_c(x_i, \theta_j) represents the logistic/softmax probability output for class cc from network jj, and nn is the ensemble size (default n=5n = 5).

    This ensembling mechanism leverages the variance across individual network instances arising from stochastic gradient optimization and random weight initialization, reducing overall prediction error variance on small datasets.

  4. Knowl 4 — One-Dimensional Temporal Receptive Field Formulation

    equation

    For a 1D convolutional neural network with depth dd (total number of convolutional layers) where each layer i∈{1,…,d}i \in \{1, \dots, d\} has a filter length kik_i and convolution stride equal to 1, the theoretical Receptive Field (RFRF) measuring the maximum field of view over the input time dimension is given by:

    RF=1+∑i=1d(ki−1)RF = 1 + \sum_{i=1}^{d} (k_i - 1)

    Adding two layers of kernel length kk increases the receptive field by 2×(k−1)2 \times (k - 1), whereas increasing the filter length kik_i across all dd existing layers by 2 increases the receptive field by 2×d2 \times d.

    In 1D temporal networks, increasing filter length expands the receptive field much more directly than adding depth, enabling the extraction of long-range patterns without requiring excessive depth that leads to overfitting on small datasets.

  5. Knowl 5 — State-of-the-Art Accuracy on the UCR Benchmark vs. HIVE-COTE and ResNet

    empirical result

    Evaluating InceptionTime across all 85 datasets of the UCR Time Series Classification archive demonstrates state-of-the-art performance:

    • Comparison with HIVE-COTE: InceptionTime achieves competitive accuracy with HIVE-COTE (the previous leading non-deep-learning ensemble of 37 classifiers), winning on 40 datasets, tying on 6, and losing on 39 (Win/Tie/Loss: 40/6/39). A Wilcoxon signed-rank test with Holm's alpha (5%) correction yields p>0.5p > 0.5, showing no statistically significant difference in classification accuracy and placing both models in the same top clique in critical difference diagrams.

    • Comparison with ResNet(5): InceptionTime significantly outperforms ResNet(5) (an ensemble of 5 ResNet architectures) with a Win/Tie/Loss of 54/8/23 (p<0.01p < 0.01) when matching batch size (64), and 53/7/25 (p<0.01p < 0.01) when comparing against ResNet with its original batch size.

    • Comparison with Component Models: InceptionTime outperforms Proximity Forest (PF), Shapelet Transform (ST), Bag-of-SFA-Symbols (BOSS), Elastic Ensemble (EE), and Nearest Neighbor Dynamic Time Warping (NN-DTW).

  6. Knowl 6 — Computational Scalability and Training Complexity vs. HIVE-COTE

    empirical result

    InceptionTime demonstrates substantial computational advantages over HIVE-COTE:

    • Theoretical vs. Practical Complexity: While HIVE-COTE incurs a training time complexity in O(N2⋅T4)\mathcal{O}(N^2 \cdot T^4) for NN time series of length TT, InceptionTime scales linearly with time series length TT and dataset size NN using standard GPU parallel acceleration.

    • Sequence Length Scaling (InlineSkate): On the InlineSkate dataset with exponential re-sampling, InceptionTime's training time increases near-linearly with TT, running nearly two orders of magnitude faster than HIVE-COTE on long series.

    • Dataset Size Scaling (SITS): On the Satellite Image Time Series (SITS) benchmark (N≈106N \approx 10^6 series, length T=46T = 46, 24 classes), HIVE-COTE requires more than 8 days to learn from N=1,500N = 1{,}500 instances, while InceptionTime learns from 1,500 instances in 1 hour and scales to 8 million instances in 13 hours on a single GPU.

  7. Knowl 7 — Synthetic Dataset Generation for Controlled Time Series Classification

    algorithm

    To evaluate the isolated effects of time series length, receptive field, depth, filter length, and number of classes, a synthetic benchmark dataset generation procedure is defined:

    Input: Time series length TT, dataset size NN, number of classes CC
    Output: Dataset D={(Xi,Yi)}i=1N\mathcal{D} = \{(X_i, Y_i)\}_{i=1}^N of z-normalized univariate series and labels
    pattern_length = ⌊0.10×T⌋\lfloor 0.10 \times T \rfloor
    D=∅\mathcal{D} = \emptyset
    for i=1i = 1 to NN:
        Yi∼UniformDiscrete(1,C)Y_i \sim \text{UniformDiscrete}(1, C)
        Xi=sample T points uniformly from U(0.0,0.1)X_i = \text{sample } T \text{ points uniformly from } \mathcal{U}(0.0, 0.1)
        tstart=ClassPatternOffset(Yi,T)t_{\text{start}} = \text{ClassPatternOffset}(Y_i, T)
        for t=0t = 0 to pattern_length - 1:
            Xi[tstart+t]=Xi[tstart+t]+1.0X_i[t_{\text{start}} + t] = X_i[t_{\text{start}} + t] + 1.0
        Xi=Xi−mean(Xi)std(Xi)X_i = \frac{X_i - \text{mean}(X_i)}{\text{std}(X_i)}
        D=D∪{(Xi,Yi)}\mathcal{D} = \mathcal{D} \cup \{(X_i, Y_i)\}
    return D\mathcal{D}

    Each class is distinguished by the fixed temporal location of an injected step pattern of amplitude 1.0 spanning 10% of the total sequence length. The surrounding uniform noise enables generating an arbitrary number of synthetic instances per class.

  8. Knowl 8 — Filter Length Selection Across Time Series Lengths

    data/table

    Comparing InceptionTime variants with short filters (InceptionTime.8 with lengths {2,4,8}\{2, 4, 8\}), long filters (InceptionTime.64 with lengths {16,32,64}\{16, 32, 64\}), and default filters (InceptionTime with lengths {10,20,40}\{10, 20, 40\}) across 85 UCR datasets binned by sequence length demonstrates the robustness of the default filter configuration:

    Length InceptionTime.8 InceptionTime.64 InceptionTime
    <81<81 1.71 2.21 1.79
    81–250 1.89 2.11 1.42
    251–450 2.45 1.32 1.86
    451–700 2.08 1.85 1.62
    701–1000 1.50 2.60 1.80
    >1000>1000 2.14 2.00 1.71

    Numbers represent the average rank (lower is better; bold indicates best) across datasets within each length range (~15 datasets per bin). The default filter setting achieves the best or second-best average rank across virtually all sequence length ranges, balancing pattern coverage against the risk of parameter overfitting on small sample sizes.

  9. Knowl 9 — Ablation of Bottleneck Layers and Residual Shortcut Connections

    empirical result

    Ablation experiments on the 85 UCR datasets evaluate the individual contributions of InceptionTime's architectural mechanisms:

    • Bottleneck Layer Ablation: InceptionTime with bottleneck layers vs. without bottleneck layers yields a Win/Tie/Loss of 39/17/29 (p>0.1p > 0.1, Wilcoxon signed-rank test). The bottleneck layer cuts the total parameter count in half without a statistically significant reduction in accuracy, balancing accuracy and computational complexity.

    • Residual Connections Ablation: InceptionTime with residual shortcuts vs. without residuals yields a Win/Tie/Loss of 38/20/27 (p>0.2p > 0.2). While overall archive accuracy is statistically indistinguishable, residual connections improve optimization convergence and prevent severe overfitting on specific complex datasets containing variable-length shapelets (such as ShapeletSim).

  10. Knowl 10 — Hyperparameter Sensitivity and Ensemble Size Ablation

    empirical result

    Ablation and sensitivity studies characterize the stability and scaling behavior of InceptionTime:

    • Ensemble Size Scaling: Testing ensemble sizes x∈{1,2,5,10,20,30}x \in \{1, 2, 5, 10, 20, 30\} shows substantial accuracy gains when increasing from a single network (x=1x = 1) to x=5x = 5. Performance gains plateau for x≥5x \ge 5, establishing 5 models as the computational sweet spot.

    • Hyperparameter Sensitivity: An alternative model, InceptionTime(second best), constructed using the second-best hyperparameter configuration from each individual ablation experiment (depth =9= 9, number of filters =16= 16, max filter length =32= 32, batch size =32= 32, bottleneck size =64= 64) compared against default InceptionTime yields a non-significant difference with p≈0.71p \approx 0.71 (Wilcoxon signed-rank test).

    This indicates that InceptionTime's high accuracy is robust and not an artifact of overfitting hyperparameters to the standard UCR benchmark splits.

Coverage note — None was omitted; all primary contributions, mathematical equations, architectural specifications, empirical benchmarks, and ablation findings have been captured.

References

  1. 1.Bagnall A, Lines J, Hills J, Bostrom A (2016) Time-series classification with COTE: The collective of transformation-based ensembles. In: International Conference on Data Engineering, pp 1548–1549
  2. 2.Bagnall A, Lines J, Bostrom A, Large J, Keogh E (2017) The great time series classification bake off: a review and experimental evaluation of recent algorithmic advances. Data Mining and Knowledge Discovery 31(3):606–660
  3. 3.Benavoli A, Corani G, Mangili F (2016) Should we really use post-hoc tests based on mean-ranks? Machine Learning Research 17(1):152–161
  4. 4.Brunel A, Pasquet J, Pasquet J, Rodriguez N, Comby F, Fouchez D, Chaumont M (2019) A CNN adapted to time series for the classification of Supernovae. In: Electronic Imaging
  5. 5.Cui Z, Chen W, Chen Y (2016) Multi-scale convolutional neural networks for time series classification. ArXiv 1603.06995
  6. 6.Cuturi M, Blondel M (2017) Soft-dtw: a differentiable loss function for time-series. In: International Conference on Machine Learnin, pp 894–903
  7. 7.Dau HA, Bagnall A, Kamgar K, Yeh CCM, Zhu Y, Gharghabi S, Ratanamahatana CA, Keogh E (2018) The ucr time series archive. ArXiv
  8. 8.Demšar J (2006) Statistical comparisons of classifiers over multiple data sets. Machine Learning Research 7:1–30
  9. 9.Forestier G, Petitjean F, Senin P, Despinoy F, Huaulmé A, Ismail Fawaz H, Weber J, Idoumghar L, Muller PA, Jannin P (2018) Surgical motion analysis using discriminative interpretable patterns. Artificial Intelligence in Medicine 91:3 – 11
  10. 10.Friedman M (1940) A comparison of alternative tests of significance for the problem of mm rankings. The Annals of Mathematical Statistics 11(1):86–92
  11. 11.Garcia S, Herrera F (2008) An extension on “statistical comparisons of classifiers over multiple data sets” for all pairwise comparisons. Machine learning research 9:2677–2694
  12. 12.Glorot X, Bengio Y (2010) Understanding the difficulty of training deep feedforward neural networks. In: International Conference on Artificial Intelligence and Statistics, vol 9, pp 249–256
  13. 13.Guan C, Wang X, Zhang Q, Chen R, He D, Xie X (2019) Towards a deep and unified understanding of deep neural models in NLP. In: International Conference on Machine Learning, pp 2454–2463
  14. 14.He K, Zhang X, Ren S, Sun J (2016) Deep residual learning for image recognition. In: IEEE Conference on Computer Vision and Pattern Recognition, pp 770–778
  15. 15.Hills J, Lines J, Baranauskas E, Mapp J, Bagnall A (2014) Classification of time series by shapelet transformation. Data Mining and Knowledge Discovery 28(4):851–881
  16. 16.Huang G, Liu Z, Van Der Maaten L, Weinberger KQ (2017) Densely connected convolutional networks. In: IEEE Conference on Computer Vision and Pattern Recognition, pp 4700–4708
  17. 17.Ismail Fawaz H, Forestier G, Weber J, Idoumghar L, Muller PA (2018) Transfer learning for time series classification. In: IEEE International Conference on Big Data, pp 1367–1376
  18. 18.Ismail Fawaz H, Forestier G, Weber J, Idoumghar L, Muller PA (2019a) Adversarial attacks on deep neural networks for time series classification. In: IEEE International Joint Conference on Neural Networks
  19. 19.Ismail Fawaz H, Forestier G, Weber J, Idoumghar L, Muller PA (2019b) Deep learning for time series classification: a review. Data Mining and Knowledge Discovery
  20. 20.Ismail Fawaz H, Forestier G, Weber J, Idoumghar L, Muller PA (2019c) Deep neural network ensembles for time series classification. In: IEEE International Joint Conference on Neural Networks
  21. 21.Ismail Fawaz H, Forestier G, Weber J, Petitjean F, Idoumghar L, Muller PA (2019d) Automatic alignment of surgical videos using kinematic data. In: Artificial Intelligence in Medicine, pp 104–113
  22. 22.Karimi-Bidhendi S, Munshi F, Munshi A (2018) Scalable classification of univariate and multivariate time series. In: IEEE International Conference on Big Data, pp 1598–1605
  23. 23.Kashiparekh K, Narwariya J, Malhotra P, Vig L, Shroff G (2019) Convtimenet: A pre-trained deep convolutional neural network for time series classification. In: IEEE International Joint Conference on Neural Networks
  24. 24.Keogh EJ, Pazzani MJ (2001) Derivative dynamic time warping. In: Proceedings of the 2001 SIAM International Conference on Data Mining, SIAM, pp 1–11
  25. 25.Kingma DP, Ba J (2015) Adam: A method for stochastic optimization. In: International Conference on Learning Representations
  26. 26.Krizhevsky A, Sutskever I, Hinton GE (2012) ImageNet Classification with Deep Convolutional Neural Networks. In: Advances in Neural Information Processing Systems, pp 1097–1105
  27. 27.Le Guennec A, Malinowski S, Tavenard R (2016) Data augmentation for time series classification using convolutional neural networks. In: ECML/PKDD Workshop on Advanced Analytics and Learning on Temporal Data
  28. 28.LeCun Y, Bottou L, Orr GB, Müller KR (1998) Efficient backprop. In: Neural Networks: Tricks of the Trade, This Book is an Outgrowth of a 1996 NIPS Workshop, pp 9–50
  29. 29.LeCun Y, Bengio Y, Hinton G (2015) Deep learning. Nature 521:436–444
  30. 30.Lee W, Park S, Joo W, Moon IC (2018) Diagnosis prediction via medical context attention networks using deep generative modeling. In: IEEE International Conference on Data Mining, pp 1104–1109
  31. 31.Lines J, Bagnall A (2015) Time series classification with ensembles of elastic distance measures. Data Mining and Knowledge Discovery 29(3):565–592
  32. 32.Lines J, Taylor S, Bagnall A (2016) HIVE-COTE: The hierarchical vote collective of transformation-based ensembles for time series classification. In: IEEE International Conference on Data Mining, pp 1041–1046
  33. 33.Liu Y, Yu J, Han Y (2018) Understanding the effective receptive field in semantic image segmentation. Multimedia Tools and Applications 77(17):22159–22171
  34. 34.Lucas B, Shifaz A, Pelletier C, O’Neill L, Zaidi N, Goethals B, Petitjean F, Webb GI (2019) Proximity forest: an effective and scalable distance-based classifier for time series. Data Mining and Knowledge Discovery 33(3):607–635
  35. 35.Luo W, Li Y, Urtasun R, Zemel R (2016) Understanding the effective receptive field in deep convolutional neural networks. In: Advances in Neural Information Processing Systems, pp 4898–4906
  36. 36.Marteau P (2009) Time warp edit distance with stiffness adjustment for time series matching. IEEE Transactions on Pattern Analysis and Machine Intelligence 31(2):306–318
  37. 37.Pelletier C, Webb GI, Petitjean F (2019) Temporal convolutional neural network for the classification of satellite image time series. Remote Sensing 11(5):523
  38. 38.Russakovsky O, Deng J, Su H, Krause J, Satheesh S, Ma S, Huang Z, Karpathy A, Khosla A, Bernstein M, Berg AC, Fei-Fei L (2015) ImageNet large scale visual recognition challenge. International Journal of Computer Vision 115(3):211–252
  39. 39.Sabour S, Frosst N, Hinton GE (2017) Dynamic routing between capsules. In: Advances in Neural Information Processing Systems, pp 3856–3866
  40. 40.Scardapane S, Wang D (2017) Randomness in neural networks: an overview. Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery 7(2):e1200
  41. 41.Schäfer P (2015a) The boss is concerned with time series classification in the presence of noise. Data Mining and Knowledge Discovery 29(6):1505–1530
  42. 42.Schäfer P (2015b) Scalable time series classification. Data Mining and Knowledge Discovery pp 1–26
  43. 43.Schäfer P, Leser U (2017) Fast and accurate time series classification with WEASEL. In: Proceedings of the 2017 ACM on Conference on Information and Knowledge Management, ACM, pp 637–646
  44. 44.Stefan A, Athitsos V, Das G (2013) The move-split-merge metric for time series. IEEE Transactions on Knowledge and Data Engineering 25(6):1425–1438
  45. 45.Szegedy C, Liu W, Jia Y, Sermanet P, Reed S, Anguelov D, Erhan D, Vanhoucke V, Rabinovich A (2015) Going deeper with convolutions. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 1–9
  46. 46.Szegedy C, Ioffe S, Vanhoucke V, Alemi AA (2017) Inception-v4, inception-resnet and the impact of residual connections on learning. In: AAAI Conference on Artificial Intelligence
  47. 47.Tan CW, Webb GI, Petitjean F (2017) Indexing and classifying gigabytes of time series under time warping. In: Proceedings of the 2017 SIAM International Conference on Data Mining, SIAM, pp 282–290
  48. 48.Vlachos M, Hadjieleftheriou M, Gunopulos D, Keogh E (2006) Indexing multidimensional time-series. The VLDB Journal—The International Journal on Very Large Data Bases 15(1):1–20
  49. 49.Wang Z, Yan W, Oates T (2017) Time series classification from scratch with deep neural networks: A strong baseline. In: International Joint Conference on Neural Networks, pp 1578–1585
  50. 50.Yi F, Yu Z, Zhuang F, Zhang X, Xiong H (2018) An integrated model for crime prediction using temporal and spatial factors. In: IEEE International Conference on Data Mining, pp 1386–1391
  51. 51.Yuan Y, Xun G, Ma F, Wang Y, Du N, Jia K, Su L, Zhang A (2018) Muvan: A multi-view attention network for multivariate temporal data. In: IEEE International Conference on Data Mining, pp 717–726
  52. 52.Zhang C, Tavanapong W, Kijkul G, Wong J, de Groen PC, Oh J (2018) Similarity-based active learning for image classification under class imbalance. In: IEEE International Conference on Data Mining, pp 1422–1427

Citation

MLA
Ismail Fawaz, H., et al. “InceptionTime: Finding AlexNet for Time Series Classification”. Data Mining and Knowledge Discovery, vol. 34, no. 6, 2020, pp. 1936–62, https://doi.org/10.1007/s10618-020-00710-y.
APA
Ismail Fawaz, H., Lucas, B., Forestier, G., Pelletier, C., Schmidt, D. F., Weber, J., Webb, G. I., Idoumghar, L., Muller, P.-A., & Petitjean, F. (2020). InceptionTime: Finding AlexNet for time series classification. Data Mining and Knowledge Discovery, 34(6), 1936–1962. https://doi.org/10.1007/s10618-020-00710-y
Chicago
Ismail Fawaz, H., B. Lucas, G. Forestier, et al. 2020. “InceptionTime: Finding AlexNet for Time Series Classification”. Data Mining and Knowledge Discovery 34 (6): 1936–62. https://doi.org/10.1007/s10618-020-00710-y.
Harvard
Ismail Fawaz, H. et al. (2020) “InceptionTime: Finding AlexNet for time series classification”, Data Mining and Knowledge Discovery, 34(6), pp. 1936–1962. Available at: https://doi.org/10.1007/s10618-020-00710-y.
Vancouver
1. Ismail Fawaz H, Lucas B, Forestier G, Pelletier C, Schmidt DF, Weber J, Webb GI, Idoumghar L, Muller P-A, Petitjean F (2020) InceptionTime: Finding AlexNet for time series classification. Data Mining and Knowledge Discovery 34:1936–1962

BibTeX

@article{Ismail_Fawaz_2020, title={InceptionTime: Finding AlexNet for time series classification}, volume={34}, ISSN={1573-756X}, url={http://dx.doi.org/10.1007/s10618-020-00710-y}, DOI={10.1007/s10618-020-00710-y}, number={6}, journal={Data Mining and Knowledge Discovery}, publisher={Springer Science and Business Media LLC}, author={Ismail Fawaz, Hassan and Lucas, Benjamin and Forestier, Germain and Pelletier, Charlotte and Schmidt, Daniel F. and Weber, Jonathan and Webb, Geoffrey I. and Idoumghar, Lhassane and Muller, Pierre-Alain and Petitjean, François}, year={2020}, month=Sept, pages={1936–1962} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF