Deep Convolutional Neural Networks and Data Augmentation for Environmental Sound Classification

Justin SalamonJuan Pablo Bello

article2016IEEE Signal Processing Letters1,479 citations2020 IEEE SPS Signal Processing Letters Best Paper Award

Demonstrates that combining deep convolutional neural networks with systematic audio data augmentation overcomes data scarcity to achieve state-of-the-art accuracy in environmental sound classification.

Listen

Automatic environmental sound classification is increasingly important for applications such as smart acoustic sensor networks for urban noise mitigation, surveillance, and context-aware computing. Deep convolutional neural networks are well-suited to identify patterns in sound spectrograms, even when interfering noise masks acoustic signals. However, deep neural networks require large quantities of labeled training data, and the relative scarcity of annotated environmental audio has previously prevented these high-capacity models from outperforming simpler, traditional machine learning techniques.

The main objective of the article is to demonstrate an effective deep convolutional neural network architecture for environmental sound classification and to evaluate how audio data augmentation—generating synthetic training variations from existing data—can overcome data scarcity to improve classification accuracy.

The authors designed a five-layer deep neural network utilizing small, localized receptive fields to capture detailed time-frequency signatures from audio spectrograms. To expand the training data without altering the semantic meaning of the sounds, they applied four audio deformations: time stretching, pitch shifting, dynamic range compression, and background noise mixing. The approach was evaluated on the UrbanSound8K dataset, which contains 8,732 real-world audio clips across ten urban sound categories, using a standardized ten-fold cross-validation methodology to benchmark against existing methods.

The analysis yielded several key findings. First, training the proposed neural network on unaugmented data resulted in an average accuracy of 73%, matching existing shallow dictionary learning models (74%) and earlier neural networks (73%). Second, applying data augmentation increased the proposed model's mean accuracy to 79%, achieving a state-of-the-art result that statistically outperformed the shallow approach. Third, expanding the capacity of the shallow dictionary model did not improve its performance, confirming that peak accuracy requires pairing a high-capacity deep network with an augmented dataset. Fourth, the impact of deformations varied significantly by sound class: pitch shifting consistently improved accuracy across all classes, while background noise and dynamic range compression harmed the classification of continuous humming sounds, such as air conditioners.

These findings indicate that deep learning can significantly improve acoustic monitoring systems, provided that data scarcity is addressed systematically. Data augmentation provides a cost-effective alternative to expensive manual data collection and labeling campaigns. However, applying indiscriminate augmentations introduces trade-offs, as certain transformations increase confusion between specific sound categories, such as confusing air conditioners with idling engines.

Based on these results, organizations deploying environmental sound recognition should pair high-capacity deep learning models with data augmentation strategies rather than relying on shallow architectures. System developers should implement class-conditional data augmentation, using validation data to apply only the specific deformations that benefit each sound category rather than applying all transformations uniformly. Further work should focus on testing these selective augmentation pipelines on broader, real-world acoustic sensor networks.

The findings are supported by a rigorous ten-fold cross-validation on a standardized urban sound benchmark. However, the study is limited to ten predefined urban sound classes in short clips of up to four seconds. Practitioners should exercise caution when deploying these specific models in environments with different acoustic characteristics or overlapping sound events not represented in the evaluation data.

arXiv: 1608.04363
Cover for Deep Convolutional Neural Networks and Data Augmentation for Environmental Sound Classification

Abstract

The ability of deep convolutional neural networks (CNN) to learn discriminative spectro-temporal patterns makes them well suited to environmental sound classification. However, the relative scarcity of labeled data has impeded the exploitation of this family of high-capacity models. This study has two primary contributions: first, we propose a deep convolutional neural network architecture for environmental sound classification. Second, we propose the use of audio data augmentation for overcoming the problem of data scarcity and explore the influence of different augmentations on the performance of the proposed CNN architecture. Combined with data augmentation, the proposed model produces state-of-the-art results for environmental sound classification. We show that the improved performance stems from the combination of a deep, high-capacity model and an augmented training set: this combination outperforms both the proposed CNN without augmentation and a "shallow" dictionary learning model with augmentation. Finally, we examine the influence of each augmentation on the model's classification accuracy for each class, and observe that the accuracy for each class is influenced differently by each augmentation, suggesting that the performance of the model could be improved further by applying class-conditional data augmentation.

Table of Contents

  • I Introduction
  • II Method
  • II-A Deep Convolutional Neural Network
  • II-B Data Augmentation
  • II-C Evaluation
  • III Results
  • IV Conclusion
  • References

Knowls

  1. Knowl 1 — SB-CNN Architecture for Environmental Sound Classification

    model/method

    The SB-CNN model is a 5-layer deep convolutional neural network designed for environmental sound classification from time-frequency audio representations.

    Input Representation: The input X∈R128×128X \in \mathbb{R}^{128 \times 128} is a time-frequency patch (TF-patch) of 128 frequency bands and 128 time frames (corresponding to 3 seconds of audio). It is extracted from a log-scaled mel-spectrogram computed with 128 mel frequency bands covering the audible range (0–22050 Hz0\text{--}22050\text{ Hz}), a window size of 23 ms23\text{ ms} (1024 samples at 44.1 kHz44.1\text{ kHz}), and a hop size of 23 ms23\text{ ms}.

    Layer Specifications:

    1. Layer ℓ1\ell_1 (Convolution + Pooling): 24 filters with receptive field (5,5)(5,5) (weight tensor shape (24,1,5,5)(24, 1, 5, 5)), followed by strided max-pooling with stride and pool shape (4,2)(4,2) over time and frequency dimensions, and rectified linear unit (ReLU) activation h(x)=max⁡(x,0)h(x) = \max(x, 0).
    2. Layer ℓ2\ell_2 (Convolution + Pooling): 48 filters with receptive field (5,5)(5,5) (weight tensor shape (48,24,5,5)(48, 24, 5, 5)), followed by (4,2)(4,2) strided max-pooling and ReLU activation.
    3. Layer ℓ3\ell_3 (Convolution): 48 filters with receptive field (5,5)(5,5) (weight tensor shape (48,48,5,5)(48, 48, 5, 5)), followed by ReLU activation with no pooling.
    4. Layer ℓ4\ell_4 (Fully Connected): 64 hidden units (weight matrix shape (2400,64)(2400, 64) after flattening layer ℓ3\ell_3 feature maps), followed by ReLU activation.
    5. Layer ℓ5\ell_5 (Output Dense): 10 output units (weight matrix shape (64,10)(64, 10)), followed by a softmax activation function: P(y=c∣X)=exp⁡(zc)∑k=110exp⁡(zk)P(y = c \mid X) = \frac{\exp(z_c)}{\sum_{k=1}^{10} \exp(z_k)} where zcz_c is the unnormalized logit for class c∈{1,…,10}c \in \{1, \dots, 10\}.

    The small (5,5)(5,5) convolutional filter receptive fields enable the network to capture localized spectro-temporal patterns that can be combined across layers into signatures resilient to partial masking by background acoustic sources.

  2. Knowl 2 — Audio Data Augmentation Deformations for Environmental Sounds

    model/method

    To overcome labeled data scarcity in environmental sound classification, semantic-preserving audio deformations are applied directly to the raw audio signals prior to log-mel-spectrogram extraction. Five augmentation sets are generated across four deformation categories:

    1. Time Stretching (TS): Modifies signal playback speed while preserving pitch. Each sample is time stretched by 4 factors: {0.81,0.93,1.07,1.23}\{0.81, 0.93, 1.07, 1.23\}
    2. Pitch Shifting Set 1 (PS1): Modifies signal pitch while preserving duration by 4 small semitone shifts: {−2,−1,+1,+2} semitones\{-2, -1, +1, +2\}\text{ semitones}
    3. Pitch Shifting Set 2 (PS2): Modifies signal pitch while preserving duration by 4 larger semitone shifts: {−3.5,−2.5,+2.5,+3.5} semitones\{-3.5, -2.5, +2.5, +3.5\}\text{ semitones}
    4. Dynamic Range Compression (DRC): Compresses the dynamic range using 4 parameter settings: 3 from the Dolby E standard (music standard, film standard, speech) and 1 from the Icecast streaming server (radio).
    5. Background Noise Mixing (BG): Blends the original audio clip xx with a background acoustic scene recording yy devoid of target classes according to: z=(1−w)⋅x+w⋅yz = (1 - w) \cdot x + w \cdot y where yy is selected from 4 acoustic scenes (street-workers, street-traffic, street-people, park), and the mixing weight ww is sampled randomly from the uniform distribution w∼U(0.1,0.5)w \sim \mathcal{U}(0.1, 0.5).
  3. Knowl 3 — SB-CNN Training and Frame-Aggregated Patch Inference Procedure

    algorithm

    The training and inference procedures for SB-CNN operate on fixed 128×128128 \times 128 time-frequency patches extracted from log-mel spectrograms.

    Training Protocol:

    • Optimization: Mini-batch stochastic gradient descent (SGD) minimizing categorical cross-entropy loss with a constant learning rate of 0.010.01.
    • Batching: Batches consist of 100 TF-patches randomly sampled without replacement. Each 3-second patch is extracted from a random temporal offset within a training clip's log-mel-spectrogram.
    • Regularization: Dropout with drop probability p=0.5p = 0.5 is applied to the inputs of dense layers ℓ4\ell_4 and ℓ5\ell_5. L2L_2 weight regularization with penalty coefficient λ=0.001\lambda = 0.001 is applied to the weights of layers ℓ4\ell_4 and ℓ5\ell_5.
    • Epoch Definition & Schedule: The model is trained for 50 epochs. An epoch finishes when minibatches exhaust 1/81/8 of all possible TF-patches (extracted from every training sample across all valid frame offsets). The model is checkpointed after each epoch.

    Validation and Test Inference Algorithm:

    Input: Audio spectrogram S of dimension 128 bands by T frames (where T >= 128), model parameters theta
    Output: Predicted class label y_hat in {1, ..., 10}
    Extract all overlapping 128-frame patches with 1-frame hop:
    P = [S[:, t : t + 128] for t in 0, 1, ..., T - 128]
    Initialize cumulative class probability vector:
    A_sum = 0 (vector of length 10)
    for each patch X_k in P:
        Compute softmax output vector: p_k = SB_CNN(X_k; theta)
        A_sum = A_sum + p_k
    Compute mean activation across all patches:
    A_mean = A_sum / |P|
    Select class with maximum mean activation:
    y_hat = argmax_c (A_mean[c])
    return y_hat

    The epoch checkpoint that achieves the highest sample-level classification accuracy on the validation split is selected for test evaluation.

  4. Knowl 4 — Classification Performance Comparison on UrbanSound8K with Data Augmentation

    empirical result

    The SB-CNN model was evaluated on the UrbanSound8K dataset (8,732 sound clips ≤4 s\le 4\text{ s} across 10 environmental sound classes) using 10-fold cross-validation, comparing performance against spherical kk-means dictionary learning (SKM) and Piczak's CNN (PiczakCNN).

    Classification Accuracy Results:

    • Without Data Augmentation:
      • SKM (k=2000k = 2000): mean accuracy of 0.740.74.
      • PiczakCNN: mean accuracy of 0.730.73.
      • SB-CNN: mean accuracy of 0.730.73.
    • With Data Augmentation (all 5 augmentation sets combined):
      • SKM (k=2000k = 2000): mean accuracy remained comparable to the unaugmented baseline. Increasing dictionary size from k=2000k = 2000 to k=4000k = 4000 yielded no further improvement.
      • SB-CNN: mean accuracy increased from 0.730.73 to 0.790.79, achieving state-of-the-art performance. The improvement over augmented SKM is statistically significant (p=0.0003p = 0.0003, paired two-sided tt-test).

    Per-Class Classification Accuracies for SB-CNN (Augmented):

    • Air conditioner: 0.490.49
    • Car horn: 0.900.90
    • Children playing: 0.830.83
    • Dog bark: 0.900.90
    • Drilling: 0.800.80
    • Engine idling: 0.800.80
    • Gun shot: 0.940.94
    • Jackhammer: 0.680.68
    • Siren: 0.850.85
    • Street music: 0.840.84

    These results show that the performance improvement requires combining a high-capacity deep convolutional model with an augmented training dataset; augmentation alone on a shallow model does not produce the same gain.

  5. Knowl 5 — Confusion Matrix and Error Distribution Shifts Under Data Augmentation

    data/table

    The confusion matrix for SB-CNN trained with all data augmentations on UrbanSound8K across 10-fold cross-validation, along with the difference matrix relative to unaugmented SB-CNN (Δ=Augmented−Unaugmented\Delta = \text{Augmented} - \text{Unaugmented}), demonstrates per-class improvements and class-pair trade-offs.

    True \ redicted AI CA CH DO DR EN GU JA SI ST
    Air Conditioner (AI) 489 10 35 52 85 262 2 25 15 25
    Car Horn (CA) 4 379 2 1 4 6 3 16 1 13
    Children Playing (CH) 17 1 830 36 18 24 6 4 13 51
    Dog Bark (DO) 12 3 34 904 5 5 4 3 15 15
    Drilling (DR) 38 0 16 39 802 8 2 69 3 23
    Engine Idling (EN) 62 7 5 8 37 798 5 60 5 13
    Gun Shot (GU) 0 0 1 19 1 0 352 0 0 1
    Jackhammer (JA) 26 1 1 0 175 48 1 673 61 14
    Siren (SI) 6 1 38 28 26 15 0 1 797 17
    Street Music (ST) 24 5 54 21 4 21 0 13 14 844
    True \ redicted (Δ\Delta) AI CA CH DO DR EN GU JA SI ST
    Air Conditioner (AI) +78 -1 -17 -40 -79 +141 +1 -4 -25 -54
    Car Horn (CA) -1 +34 -8 -6 -4 0 +3 +7 +1 -26
    Children Playing (CH) +5 +1 +23 -31 -3 +6 +5 0 +2 -8
    Dog Bark (DO) -8 -6 -13 +37 -5 +2 -2 +3 +1 -9
    Drilling (DR) +15 -3 -11 -8 +40 -19 -2 +15 -26 -1
    Engine Idling (EN) +3 0 -34 -29 +15 +132 +2 -34 -45 -10
    Gun Shot (GU) -1 0 +1 -9 -3 -1 +14 -1 0 0
    Jackhammer (JA) -37 +1 -17 -2 -23 -48 0 +141 0 -15
    Siren (SI) 0 +1 -1 -25 +14 +6 0 0 +16 -11
    Street Music (ST) +7 0 -39 -3 -5 +13 0 +6 -6 +27

    Key trends:

    • Diagonal values of Δ\Delta are positive for all 10 classes (ranging from +14+14 for gun shots to +141+141 for jackhammers and +132+132 for engine idling), showing universal true-positive improvements across classes.
    • Off-diagonal shifts reveal class-pair trade-offs: confusion between air conditioner and drilling decreased substantially (−79-79), but confusion between air conditioner and engine idling increased (+141+141).
  6. Knowl 6 — Class-Specific Sensitivity to Audio Augmentation Deformations

    empirical result

    Evaluating individual augmentation sets (Time Stretching [TS], Pitch Shifting [PS1 and PS2], Dynamic Range Compression [DRC], Background Noise [BG], and all combined [All]) on SB-CNN reveals heterogeneous effects across sound classes:

    • Pitch Shifting Dominance: Pitch shifting (both PS1 and PS2) yields the largest positive accuracy gains and is the only augmentation family that does not decrease accuracy for any of the 10 classes.
    • Degradation on Continuous Stationary Sounds: Dynamic range compression (DRC) and background noise (BG) decrease classification accuracy for the air conditioner class. Because air conditioner audio is characterized by a continuous, low-intensity background "hum", additive noise masks distinguishing spectral features.
    • Sub-optimality of Global Augmentation: Only 5 of the 10 sound classes achieve their highest classification accuracy when all augmentations are combined compared to using an optimal class-specific subset. This indicates that applying class-conditional data augmentation based on validation split tuning can further improve overall classification performance.

Coverage note — None was omitted; all contributed methods, model architecture parameters, data augmentation designs, experimental results, and class sensitivity analyses are included.

References

  1. 1.S. Chu, S. Narayanan, and C.-C. Kuo, “Environmental sound recognition with time-frequency audio features,” IEEE Trans. on Audio, Speech, and Language Processing, vol. 17, no. 6, pp. 1142–1158, Aug. 2009.
  2. 2.R. Radhakrishnan, A. Divakaran, and P. Smaragdis, “Audio analysis for surveillance applications,” in IEEE Worksh. on Apps. of Signal Processing to Audio and Acoustics (WASPAA’05), New Paltz, NY, USA, Oct. 2005, pp. 158–161.
  3. 3.C. Mydlarz, J. Salamon, and J. P. Bello, “The implementation of low-cost urban acoustic monitoring devices,” Applied Acoustics, vol. In Press, 2016.
  4. 4.A. Mesaros, T. Heittola, O. Dikmen, and T. Virtanen, “Sound event detection in real life recordings using coupled matrix factorization of spectral representations and class activity annotations,” in International Conference on Acoustics, Speech and Signal Processing (ICASSP), Brisbane, Australia, Apr. 2015, pp. 151–155.
  5. 5.E. Benetos, G. Lafay, M. Lagrange, and M. D. Plumbley, “Detection of overlapping acoustic events using a temporally-constrained probabilistic model,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Shanghai, China, Mar. 2016, pp. 6450–6454.
  6. 6.V. Bisot, R. Serizel, S. Essid, and G. Richard, “Acoustic scene classification with matrix factorization for unsupervised feature learning,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Shanghai, China, Mar. 2016, pp. 6445–6449.
  7. 7.J. Salamon and J. P. Bello, “Unsupervised feature learning for urban sound classification,” in IEEE Int. Conf. on Acoustics, Speech and Signal Processing (ICASSP), Brisbane, Australia, Apr. 2015, pp. 171–175.
  8. 8.——, “Feature learning with deep scattering for urban sound analysis,” in 2015 European Signal Processing Conference, Nice, France, Aug. 2015.
  9. 9.J. T. Geiger and K. Helwani, “Improving event detection for audio surveillance using gabor filterbank features,” in 23rd European Signal Processing Conference (EUSIPCO), Nice, France, Aug. 2015, pp. 714–718.
  10. 10.E. Cakir, T. Heittola, H. Huttunen, and T. Virtanen, “Polyphonic sound event detection using multi label deep neural networks,” in 2015 International Joint Conference on Neural Networks (IJCNN), July 2015, pp. 1–7.
  11. 11.K. J. Piczak, “Environmental sound classification with convolutional neural networks,” in 25th International Workshop on Machine Learning for Signal Processing (MLSP), Boston, MA, USA, Sep. 2015, pp. 1–6.
  12. 12.D. Giannoulis, E. Benetos, D. Stowell, M. Rossignol, M. Lagrange, and M. D. Plumbley, “Detection and classification of acoustic scenes and events: An IEEE AASP challenge,” in IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA), New Paltz, NY, USA, Oct. 2013, pp. 1–4.
  13. 13.D. Stowell, D. Giannoulis, E. Benetos, M. Lagrange, and M. D. Plumbley, “Detection and classification of acoustic scenes and events,” IEEE Transactions on Multimedia, vol. 17, no. 10, pp. 1733–1746, Oct. 2015.
  14. 14.S. Sigtia, A. Stark, S. Krstulovic, and M. Plumbley, “Automatic environmental sound recognition: Performance versus computational cost,” IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. PP, no. 99, pp. 1–1, 2016.
  15. 15.Y. Lecun, L. Bottou, Y. Bengio, and P. Haffner, “Gradient-based learning applied to document recognition,” Proceedings of the IEEE, vol. 86, no. 11, pp. 2278–2324, Nov. 1998.
  16. 16.C. V. Cotton and D. P. W. Ellis, “Spectral vs. spectro-temporal features for acoustic event detection,” in IEEE Worksh. on Apps. of Signal Processing to Audio and Acoustics (WASPAA), New Paltz, NY, USA, Oct. 2011, pp. 69–72.
  17. 17.J. Salamon, C. Jacoby, and J. P. Bello, “A dataset and taxonomy for urban sound research,” in 22nd ACM International Conference on Multimedia (ACM-MM’14), Orlando, FL, USA, Nov. 2014, pp. 1041–1044.
  18. 18.K. J. Piczak, “ESC: Dataset for environmental sound classification,” in 23rd ACM International Conference on Multimedia, Brisbane, Australia, Oct. 2015, pp. 1015–1018.
  19. 19.A. Mesaros, E. Fagerlund, A. Hiltunen, T. Heittola, and T. Virtanen, “TUT sound events 2016, development dataset,” Available online: http://dx.doi.org/10.5281/zenodo.45759 (accessed 10 August 2016), 2016. [Online]. Available: http://dx.doi.org/10.5281/zenodo.45759
  20. 20.A. Krizhevsky, I. Sutskever, and G. Hinton, “ImageNet classification with deep convolutional neural networks,” in Advances in neural information processing systems (NIPS), 2012, pp. 1097–1105.
  21. 21.P. Y. Simard, D. Steinkraus, and J. C. Platt, “Best practices for convolutional neural networks applied to visual document analysis.” in International Conference on Document Analysis and Recognition, vol. 3, Edinburgh, Scottland, UK, Aug. 2003, pp. 958–962.
  22. 22.B. McFee, E. Humphrey, and J. Bello, “A software framework for musical data augmentation,” in 16th Int. Soc. for Music Info. Retrieval Conf., Malaga, Spain, Oct. 2015, pp. 248–254.
  23. 23.G. Parascandolo, H. Huttunen, and T. Virtanen, “Recurrent neural networks for polyphonic sound event detection in real life recordings,” in International Conference on Acoustics, Speech and Signal Processing (ICASSP), Shanghai, China, Mar. 2016, pp. 6440–6444.
  24. 24.D. Bogdanov, N. Wack, E. Gomez, S. Gulati, P. Herrera, O. Mayor, G. Roma, J. Salamon, J. Zapata, and X. Serra, “ESSENTIA: an audio analysis library for music information retrieval,” in 14th Int. Soc. for Music Info. Retrieval Conf., Curitiba, Brazil, Nov. 2013, pp. 493–498.
  25. 25.L. Bottou, “Large-scale machine learning with stochastic gradient descent,” in 19th International Conference on Computational Statistics (COMPSTAT), Paris, France, Aug. 2010, pp. 177–186. [Online]. Available: http://dx.doi.org/10.1007/978-3-7908-2604-3 16
  26. 26.N. Srivastava, G. E. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov, “Dropout: a simple way to prevent neural networks from overfitting,” Journal of Machine Learning Research, vol. 15, no. 1, pp. 1929–1958, 2014.
  27. 27.S. Dieleman, J. Schluter, C. Raffel, E. Olson, S. Sønderby, D. Nouri, D. Maturana, M. Thoma, E. Battenberg, and J. Kelly, “Lasagne: First release,” https://github.com/Lasagne/Lasagne, 2015. [Online]. Available: http://dx.doi.org/10.5281/zenodo.27878
  28. 28.B. McFee and E. J. Humphrey, “pescador: 0.1.0,” https://github.com/bmcfee/pescador, 2015. [Online]. Available: http://dx.doi.org/10.5281/zenodo.32468
  29. 29.Dolby Labortories, Inc., “Standards and practices for authoring Dolby Digital and Dolby E bitstreams,” 2002.
  30. 30.“Icecast streaming media server forum,” (accessed 12 August 2016). [Online]. Available: http://icecast.imux.net/viewtopic.php?t=3462
  31. 31.E. J. Humphrey, J. Salamon, O. Nieto, J. Forsyth, R. Bittner, and J. P. Bello, “JAMS: A JSON annotated music specification for reproducible MIR research,” in 15th Int. Soc. for Music Info. Retrieval Conf., Taipei, Taiwan, Oct. 2014, pp. 591–596.
  32. 32.B. McFee, E. J. Humphrey, O. Nieto, J. Salamon, R. Bittner, J. Forsyth, and J. P. Bello, “Pump up the JAMS: V0.2 and beyond,” Music and Audio Research Laboratory, New York University, Tech. Rep., Oct. 2015.

Citation

MLA
Salamon, J., and J. P. Bello. “Deep Convolutional Neural Networks and Data Augmentation for Environmental Sound Classification”. IEEE Signal Processing Letters, vol. 24, no. 3, 2017, pp. 279–83, https://doi.org/10.1109/LSP.2017.2657381.
APA
Salamon, J., & Bello, J. P. (2017). Deep Convolutional Neural Networks and Data Augmentation for Environmental Sound Classification. IEEE Signal Processing Letters, 24(3), 279–283. https://doi.org/10.1109/LSP.2017.2657381
Chicago
Salamon, J., and J. P. Bello. 2017. “Deep Convolutional Neural Networks and Data Augmentation for Environmental Sound Classification”. IEEE Signal Processing Letters 24 (3): 279–83. https://doi.org/10.1109/LSP.2017.2657381.
Harvard
Salamon, J. and Bello, J.P. (2017) “Deep Convolutional Neural Networks and Data Augmentation for Environmental Sound Classification”, IEEE Signal Processing Letters, 24(3), pp. 279–283. Available at: https://doi.org/10.1109/LSP.2017.2657381.
Vancouver
1. Salamon J, Bello JP (2017) Deep Convolutional Neural Networks and Data Augmentation for Environmental Sound Classification. IEEE Signal Processing Letters 24:279–283

BibTeX

@article{Salamon_2017, title={Deep Convolutional Neural Networks and Data Augmentation for Environmental Sound Classification}, volume={24}, ISSN={1558-2361}, url={http://dx.doi.org/10.1109/LSP.2017.2657381}, DOI={10.1109/lsp.2017.2657381}, number={3}, journal={IEEE Signal Processing Letters}, publisher={Institute of Electrical and Electronics Engineers (IEEE)}, author={Salamon, Justin and Bello, Juan Pablo}, year={2017}, month=Mar, pages={279–283} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF