Open-Source Conversational AI with SpeechBrain 1.0

Mirco RavanelliTitouan ParcolletAdel MoumenSylvain de LangenCem SubakanPeter PlantingaYingzhi WangPooneh MousaviLuca Della LiberaArtem Ploujnikov

article2024JMLR107 citations

Presents SpeechBrain 1.0, an open-source PyTorch framework featuring over 200 reproducible training recipes, pre-trained models, and a unified benchmark suite supporting speech, audio, language model integration, and neurotechnology tasks.

Listen

The rapid advancement of conversational artificial intelligence—spanning voice assistants and large language models—faces a growing reproducibility crisis. Many current initiatives release only pre-trained model weights while keeping critical training code, algorithms, and data undisclosed. This lack of transparency limits scientific verification, slows innovation, and introduces adoption risks for organizations that require fully auditable systems. The article demonstrates the release of SpeechBrain 1.0, an open-source, PyTorch-based conversational AI toolkit designed to resolve these challenges by providing fully reproducible models alongside complete training recipes, dataset manifests, and hyperparameter configurations.

The developers evaluated and built the toolkit across speech recognition, synthesis, audio enhancement, natural language processing, and brain-signal decoding. The project incorporates over 200 training recipes, more than 100 pre-trained models, and community benchmark datasets. The implementation prioritizes accessible data, modular architecture, and standardized interfaces to evaluate system performance and scalability across multi-GPU environments.

The findings show that fully open-source conversational AI can achieve state-of-the-art performance while maintaining complete transparency. First, independent model replications consistently matched or exceeded original benchmarks; for example, the toolkit's speaker verification model reduced the equal error rate from 0.87% to 0.81% relative to the original published baseline through improved data augmentation and hyperparameter tuning. Second, the framework demonstrated enterprise-grade scalability by training a 1-billion-parameter speech representation model on 14,000 hours of speech using over 100 graphics processing units. Third, the toolkit successfully integrates speech with text-based large language models and discrete audio tokenizers, enabling advanced tasks such as prompted rescoring, joint audio-speech understanding, and dialogue response generation. Fourth, it expands conventional speech processing to include electroencephalography (EEG) brain signals, demonstrating that shared neural network architectures can effectively process both speech and non-verbal neural inputs.

These results demonstrate that organizations do not need to rely on proprietary or opaque AI systems to achieve high performance. By utilizing end-to-end recipes where 95% of tasks rely on freely accessible data, organizations can lower licensing costs, mitigate vendor lock-in, and audit models for safety and compliance. The findings also prove that structured training pipelines reduce the engineering overhead of deploying complex speech and language systems.

Decision-makers and engineering teams looking to build speech and conversational pipelines should evaluate open-source toolkits like SpeechBrain for prototyping and production. Organizations should adopt standardized benchmarking suites to ensure fair cross-model comparison and explore discrete audio token frameworks when planning future multimodal AI deployments. Moving forward, continued engineering is planned to develop unified multimodal foundation models that natively process text, speech, and audio within a single architecture.

Confidence in the reported capabilities is high given the toolkit's broad community adoption, encompassing 2.5 million monthly downloads and extensive replication testing. However, decision-makers should note that training large-scale foundation models from scratch remains computationally resource-intensive, requiring substantial GPU infrastructure. In addition, experimental modalities such as neural signal decoding represent an early-stage research frontier compared to mature speech recognition and synthesis workflows.

  • Paper: MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark, S. Sakshi et al. (2025). Establishes a massive multi-task audio reasoning and understanding benchmark (MMAU), extending the evaluation of comprehensive speech and audio processing systems like those enabled by SpeechBrain.
  • Paper: Qwen3-TTS Technical Report, Hangrui Hu et al. (2026). Applies modern dual-track speech tokenization and generative LLM integration to advance controllable text-to-speech beyond traditional conversational AI recipes.
  • Paper: Qwen3.5-Omni Technical Report, Qwen Team (2026). Extends conversational and speech AI into a fully unified omnimodal Thinker-Talker framework capable of real-time end-to-end multimodal perception and speech synthesis.
Cover for Open-Source Conversational AI with SpeechBrain 1.0

Abstract

SpeechBrain1 is an open-source Conversational AI toolkit based on PyTorch, focused particularly on speech processing tasks such as speech recognition, speech enhancement, speaker recognition, text-to-speech, and much more. It promotes transparency and replicability by releasing both the pre-trained models and the complete “recipes” of code and algorithms required for training them. This paper presents SpeechBrain 1.0, a significant milestone in the evolution of the toolkit, which now has over 200 recipes for speech, audio, and language processing tasks, and more than 100 models available on Hugging Face. SpeechBrain 1.0 introduces new technologies to support diverse learning modalities, Large Language Model (LLM) integration, and advanced decoding strategies, along with novel models, tasks, and modalities. It also includes a new benchmark repository, offering researchers a unified platform for evaluating models across diverse tasks.

Table of Contents

  • 1. Introduction
  • 2. Overview of SpeechBrain
  • 2.1 Architecture Overview
  • 3. Recent Developments
  • 4. Conclusion and Future Work
  • Acknowledgment
  • References
  • Appendix A. Related Toolkits
  • Appendix B. Model Replication

Knowls

  1. Knowl 1 — SpeechBrain 1.0 Architecture and Training Workflow

    model/method

    SpeechBrain 1.0 is an open-source PyTorch-based conversational AI toolkit designed for speech, audio, text, and neurotechnology (EEG) processing under the Apache 2.0 license. The framework structures model development around three decoupled components:

    1. Data Manifests: Datasets are declared using standardized CSV or JSON files specifying item identifiers, audio file paths, durations, transcriptions, speaker IDs, and timestamp segments.
    2. HyperPyYAML Configuration: A custom extension of YAML that enables hierarchical definition and runtime instantiation of Python objects, models, optimizers, learning rate schedulers, evaluation metrics, and hyperparameter dictionaries.
    3. Brain Class and Training Script: A Python execution script defining a subclass of the core Brain abstraction. Users specify task-specific computation routines by overriding methods such as forward(), compute_objectives(), and compute_metrics(), while high-level execution, distributed multi-GPU training, sequence-to-sequence loops, checkpointing, and evaluation are managed by fit() and eval().

    All components rely on standard PyTorch interfaces (torch.nn.Module, torch.optim, and torch.utils.data.Dataset), facilitating native interoperability across the PyTorch ecosystem.

  2. Knowl 2 — Supported Modalities and Processing Tasks in SpeechBrain 1.0

    model/method

    SpeechBrain 1.0 natively supports deep learning tasks and processing pipelines spanning four modalities:

    • Audio: Neural vocoding, audio data augmentation, acoustic feature extraction, sound event detection (SED), and multi-channel beamforming.
    • Speech: Automatic speech recognition (ASR), speech enhancement, source separation, text-to-speech (TTS) synthesis, speaker recognition and verification, speech-to-speech translation (S2ST), spoken language understanding (SLU), voice activity detection (VAD), speaker diarization, speech emotion recognition (SER), speech emotion diarization, spoken language identification (LID), self-supervised learning (SSL) pre-training, metric learning, and forced phonetic alignment.
    • Text: Autoregressive language model (LM) training, large language model (LLM) fine-tuning, spoken dialogue modeling, conversational response generation, and grapheme-to-phoneme (G2P) transcription.
    • Electroencephalography (EEG): Brain-computer interface (BCI) decoding tasks, including motor imagery, P300 event-related potential detection, and steady-state visual evoked potential (SSVEP) classification.
  3. Knowl 3 — Learning Modalities, Interpretability, and Self-Supervised Pre-training in SpeechBrain 1.0

    model/method

    SpeechBrain 1.0 integrates several learning paradigms for audio and speech models:

    • Continual Learning: Provides rehearsal-based, architecture-based, and regularization-based algorithms to adapt speech recognition and audio models to new distributions without catastrophic forgetting.
    • Interpretability and Explainability: Includes post-hoc and design-based explainability tools such as Post-hoc Interpretation via Quantization (PIQ), Listen to Interpret using Non-negative Matrix Factorization (NMF), Activation Map Thresholding (AMT) for focal networks, and Listenable Maps for audio classifiers.
    • Diffusion-Based Generation: Supports standard and latent diffusion models for audio synthesis and incorporates DiffWave as a diffusion-based neural vocoder.
    • Large-Scale SSL Pre-training: Features from-scratch pre-training implementations of self-supervised speech representations, including wav2vec 2.0 (scaled to a 1-billion-parameter French foundation model on 14,000 hours of speech using over 100 A100 GPUs) and the BERT-based Speech pre-Training with Random Quantization (BEST-RQ) model, along with parameter-efficient fine-tuning strategies for fast downstream inference.
  4. Knowl 4 — Neural Architectures, Discrete Tokenizers, and EEG Models in SpeechBrain 1.0

    model/method

    SpeechBrain 1.0 provides implementations of specialized neural architectures across domains:

    • Acoustic and Sequence Architectures: Implements Conformer, HyperConformer (multi-head hypermixer architecture), Branchformer (parallel MLP-attention architecture capturing local and global contextual dependencies), and Streamable Conformer Transducers. It also includes Stabilised Light Gated Recurrent Units (SLi-GRU), a numerically stabilized and accelerated variant of the Light GRU for efficient recurrent sequence modeling.
    • Discrete Audio Token Extraction: Provides interfaces for extracting and quantizing audio into discrete tokens using representations from discrete wav2vec 2.0, HuBERT, WavLM, EnCodec, Descript Audio Codec (DAC), and SpeechTokenizer for multimodal language modeling.
    • Speech Emotion Diarization: Supports temporal segmentation and simultaneous tracking of emotional state transitions across speech streams.
    • EEG Architectures and Frameworks: Implements deep neural networks for electroencephalographic decoding, including EEGNet (compact depthwise-separable CNN), ShallowConvNet (shallow convolutional network for oscillatory signals), and EEGConformer (convolutional transformer combining temporal-spatial convolutions with self-attention). Data loading and preprocessing pipelines interface with MOABB, MNE, and Braindecode.
  5. Knowl 5 — Modular Decoding Strategies and Language Model Rescoring in SpeechBrain 1.0

    model/method

    SpeechBrain 1.0 refactors beam search decoding by completely decoupling search algorithms from scoring functions:

    • Modular Search and Scoring: Enables arbitrary external scorers (including statistical nn-gram language models, neural language models, and heuristic scorers) to be plugged directly into beam search pipelines.
    • ASR & Transducer Decoding: Supports pure Connectionist Temporal Classification (CTC) training and decoding, latency-controlled Recurrent Neural Network Transducer (RNN-T) beam search, and parallel batched GPU decoding.
    • External Toolkit Integration: Interfaces with Kaldi2 (k2) for Finite State Acceptor (FSA) and Finite State Transducer (FST) graph search and decoding, and integrates KenLM for fast nn-gram language model evaluation and rescoring.
    • N-Best Rescoring: Generates ranked NN-best hypothesis lists from acoustic decoders for downstream rescoring with masked or autoregressive neural language models.
  6. Knowl 6 — Large Language Model (LLM) Integration and Rescoring in SpeechBrain 1.0

    model/method

    SpeechBrain 1.0 incorporates Large Language Model (LLM) workflows for spoken conversational AI:

    • Spoken Dialogue Modeling: Features fine-tuning pipelines for autoregressive language models (such as GPT-2, Llama 2, and Llama 3) for dialogue state tracking and conversational agent response generation.
    • Joint Audio-Speech LLMs: Implements LTU-AS (Listen, Think, and Understand - Audio and Speech), enabling foundation models to jointly ingest, reason over, and transcribe mixed acoustic audio and spoken language inputs.
    • Prompted Generative Rescoring: Implements generative ASR hypothesis rescoring (such as Progres), where LLMs are prompted with NN-best candidate transcriptions to re-rank and select the most contextually accurate hypothesis.
  7. Knowl 7 — Standardized Benchmarking Suites in SpeechBrain 1.0

    model/method

    SpeechBrain 1.0 includes four standardized evaluation and benchmark repositories:

    1. CL-MASR: A benchmark for evaluating continual learning strategies (rehearsal, regularization, architectural methods) across sequential multilingual Automatic Speech Recognition tasks.
    2. MP3S: A benchmarking framework for evaluating speech self-supervised learning (SSL) representations using modular, customizable downstream probing heads.
    3. DASB (Discrete Audio and Speech Benchmark): A benchmark dedicated to assessing discrete audio representations and neural audio codecs across speech recognition, synthesis, enhancement, and classification tasks.
    4. SpeechBrain-MOABB: A benchmark integration based on the Mother of all BCI Benchmarks (MOABB) and MNE for systematically evaluating deep neural networks on EEG classification tasks.
  8. Knowl 8 — ECAPA-TDNN Speaker Verification Replication Performance

    data/table

    SpeechBrain 1.0 provides an open-source re-implementation and training recipe for the ECAPA-TDNN (Emphasized Channel Attention, Propagation and Aggregation Time Delay Neural Network) architecture for speaker verification. Compared to the baseline reported in the original paper, the SpeechBrain recipe achieves a lower Equal Error Rate (EER):

    Model Implementation Equal Error Rate (EER %)
    Original Paper 0.87
    SpeechBrain Re-implementation 0.81

    The improved performance of the SpeechBrain re-implementation is driven by a more robust audio data augmentation pipeline and refined training hyperparameter selection.

Coverage note — None was omitted; all major architectural updates, supported modalities, learning paradigms, decoding strategies, LLM integrations, benchmarks, and replication results from SpeechBrain 1.0 are covered.

References

  1. 1.B. Aristimunha, I. Carrara, P. Guetschel, S. Sedlar, P. Rodrigues, J. Sosulski, D. Narayanan, E. Bjareholt, Q. Barthelemy, R. Kobler, R. T. Schirrmeister, E. Kalunga, L. Darmet, C. Gre- goire, A. Abdul Hussain, R. Gatti, V. Goncharenko, J. Thielen, T. Moreau, Y. Roy, V. Ja- yaram, A. Barachant, and S. Chevallier. Mother of all BCI Benchmarks, 2024. URL https: //github.com/NeuroTechX/moabb.
  2. 2.A. Baevski, H. Zhou, A. Mohamed, and M. Auli. wav2vec 2.0: A framework for self-supervised learning of speech representations. In Proceedings of the International Conference on Neural Information Processing Systems (NIPS), 2020a.
  3. 3.A. Baevski, Y. Zhou, A. Mohamed, and M. Auli. wav2vec 2.0: A framework for self-supervised learning of speech representations. In Proceedings of the International Conference on Neural Information Processing Systems (NeurIPS), 2020b.
  4. 4.D. Borra, F. Paissan, and M. Ravanelli. SpeechBrain-MOABB: An open-source Python library for benchmarking deep neural networks applied to EEG signals. Computers in Biology and Medicine, 182:97–109, 2024.
  5. 5.H. Bredin. pyannote.audio 2.1 speaker diarization pipeline: principle, benchmark, and recipe. In Proceedings of Interspeech, 2023.
  6. 6.L. Della Libera, P. Mousavi, S. Zaiem, C. Subakan, and M. Ravanelli. CL-MASR: A Continual Learning Benchmark for Multilingual ASR. CoRR, abs/2310.16931, 2023.
  7. 7.L. Della Libera, C. Subakan, and M. Ravanelli. Focal modulation networks for interpretable sound classification. In Proceedings of the ICASSP Workshop on Explainable AI for Speech and Audio (XAI-SA), 2024.
  8. 8.B. Desplanques, J. Thienpondt, and K. Demuynck. ECAPA-TDNN: emphasized channel attention, propagation and aggregation in TDNN based speaker verification. In Proceedings of Interspeech, 2020.
  9. 9.Y. Gong, A. H. Liu, H. Luo, L. Karlinsky, and J. Glass. Joint audio and speech understanding. In In Proceedings of the IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), 2023.
  10. 10.A. Gramfort, M. Luessi, E. Larson, D. A. Engemann, D. Strohmeier, C. Brodbeck, L. Parkkonen, and M. S. H¨am¨al¨ainen. Mne software for processing meg and eeg data. NeuroImage, 86:446–460, 2014.
  11. 11.A. Gulati, J. Qin, C.-C. Chiu, N. Parmar, Y. Zhang, J. Yu, W. Han, S. Wang, Z. Zhang, Y. Wu, and R. Pang. Conformer: Convolution-augmented transformer for speech recognition. In Proceedings of Interspeech, 2020.
  12. 12.K. Heafield. KenLM: Faster and Smaller Language Model Queries. In Proceedings of the Sixth Workshop on Statistical Machine Translation (WMT), 2011.
  13. 13.J. Hwang, M. Hira, C. Chen, X. Zhang, Z. Ni, G. Sun, P. Ma, R. Huang, V. Pratap, Y. Zhang, A. Kumar, C.-Y. Yu, C. Zhu, C. Liu, J. Kahn, M. Ravanelli, P. Sun, S. Watanabe, Y. Shi, Y. Tao, R. Scheibler, S. Cornell, S. Kim, and S. Petridis. Torchaudio 2.1: Advancing speech recognition, self-supervised learning, and audio processing components for pytorch. In Proceedings of the IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), 2023.
  14. 14.M. Jain, K. Schubert, J. Mahadeokar, C. Yeh, K. Kalgaonkar, A. Sriram, C. Fuegen, and M. L. Seltzer. RNN-T for latency controlled ASR with improved beam search. CoRR, abs/1911.01629, 2019.
  15. 15.W. Kang, L. Guo, F. Kuang, L. Lin, M. Luo, Z. Yao, X. Yang, P. Zelasko, and D. Povey. Fast and Parallel Decoding for Transducer. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2023.
  16. 16.S. Kapoor and A. Narayanan. Leakage and the reproducibility crisis in machine-learning-based science. Patterns, 4(9), 2023.
  17. 17.S. Kim, T. Hori, and S. Watanabe. Joint ctc-attention based end-to-end speech recognition using multi-task learning. In Proceedings of the IEEE international conference on acoustics, speech and signal processing (ICASSP), pages 4835–4839, 2017.
  18. 18.J. Kong, J. Kim, and J. Bae. Hifi-gan: generative adversarial networks for efficient and high fidelity speech synthesis. In Proceedings of the International Conference on Neural Information Processing Systems (NIPS), 2020a.
  19. 19.Z. Kong, W. Ping, J. Huang, K. Zhao, and B. Catanzaro. Diffwave: A Versatile Diffusion Model for Audio Synthesis. CoRR, abs/2009.09761, 2020b.
  20. 20.O. Kuchaiev, J. Li, H. Nguyen, O. Hrinchuk, R. Leary, B. Ginsburg, S. Kriman, S. Beliaev, V. Lavrukhin, J. Cook, P. Castonguay, M. Popova, J. Huang, and J. M. Cohen. NeMo: a toolkit for building AI applications using Neural Modules. CoRR, abs/1909.09577, 2019.
  21. 21.V. J. Lawhern, A. J. Solon, N. R. Waytowich, S. M. Gordon, C. P. Hung, and B. J. Lance. EEGNet: a compact convolutional neural network for EEG-based brain computer interfaces. Journal of Neural Engineering, 15(5), July 2018.
  22. 22.C. Li, L. Yang, W. Wang, and Y. Qian. Skim: Skipping memory lstm for low-latency real-time continuous speech separation. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2022.
  23. 23.A. Liesenfeld and M. Dingemanse. Rethinking open source generative AI: open washing and the EU AI Act. In Proceedings of the ACM Conference on Fairness, Accountability, and Transparency, 2024.
  24. 24.Y. Luo and N. Mesgarani. Conv-TasNet: Surpassing Ideal Time–Frequency Magnitude Masking for Speech Separation. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 27(8): 1256–1266, aug 2019.
  25. 25.Y. Luo, Z. Chen, and T. Yoshioka. Dual-path RNN: efficient long sequence modeling for time- domain single-channel speech separation. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2020.
  26. 26.F. Mai, J. Zuluaga-Gomez, T. Parcollet, and P. Motlicek. Hyperconformer: Multi-head hypermixer for efficient speech recognition. In Proceedings of Interspeech, 2023.
  27. 27.M. McTear. Conversational AI: Dialogue Systems, Conversational Agents, and Chatbots. Synthesis lectures on human language technologies. Morgan & Claypool Publishers, 2021.
  28. 28.A. Moumen and T. Parcollet. Stabilising and accelerating light gated recurrent units for automatic speech recognition. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2023.
  29. 29.P. Mousavi, J. Duret, S. Zaiem, L. D. Libera, A. Ploujnikov, C. Subakan, and M. Ravanelli. How should we extract discrete audio tokens from self-supervised models? In Proceedings of Interspeech, 2024a.
  30. 30.P. Mousavi, L. D. Libera, J. Duret, A. Ploujnikov, C. Subakan, and M. Ravanelli. DASB-Discrete Audio and Speech Benchmark. CoRR, abs/2406.14294, 2024b.
  31. 31.S. M. Mousavi, G. Roccabruna, S. Alghisi, M. Rizzoli, M. Ravanelli, and G. Riccardi. Are LLMs Robust for Spoken Dialogues? In Proceedings of the International Workshop on Spoken Dialogue Systems Technology (IWSDS), 2024c.
  32. 32.F. Paissan, C. Subakan, and M. Ravanelli. Posthoc Interpretation via Quantization. CoRR, abs/2303.12659, 2023.
  33. 33.F. Paissan, M. Ravanelli, and C. Subakan. Listenable Maps for Audio Classifiers. In Proceedings of the International Conference on Machine Learning (ICML), 2024.
  34. 34.T. Parcollet, H. Nguyen, S. Evain, M. Zanon Boito, A. Pupier, S. Mdhaffar, H. Le, S. Alisamir, N. Tomashenko, M. Dinarelli, S. Zhang, A. Allauzen, M. Coavoux, Y. Est`eve, M. Rouvier, J. Gou- lian, B. Lecouteux, F. Portet, S. Rossato, F. Ringeval, D. Schwab, and L. Besacier. LeBenchmark 2.0: A standardized, replicable and enhanced framework for self-supervised representations of French speech. Computer Speech & Language, 86:101622, 2024.
  35. 35.J. Parekh, S. Parekh, P. Mozharovskyi, F. Alche-Buc, and G. Richard. Listen to Interpret: Post-hoc Interpretability for Audio Networks with NMF. In In proceedings of the International Conference on Neural Information Processing Systems (NeurIPS), 2022.
  36. 36.Y. Peng, S. Dalmia, I. Lane, and S. Watanabe. Branchformer: Parallel MLP-Attention Architectures to Capture Local and Global Context for Speech Recognition and Understanding. In Proceedings of the International Conference on Machine Learning (ICML), 2022a.
  37. 37.Y. Peng, S. Dalmia, I. R. Lane, and S. Watanabe. Branchformer: Parallel MLP-Attention Archi- tectures to Capture Local and Global Context for Speech Recognition and Understanding. In Proceedings of the International Conference on Machine Learning (ICML), 2022b.
  38. 38.A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever. Language models are unsu- pervised multitask learners. Technical report, OpenAI, 2019. Technical report.
  39. 39.M. Ravanelli and M. Omologo. On the selection of the impulse responses for distant-speech recog- nition based on contaminated speech training. In Proceesings of Interspeech, 2014.
  40. 40.M. Ravanelli and M. Omologo. Contaminated speech training methods for robust DNN-HMM distant speech recognition. In Proceesings of Interspeech, 2015.
  41. 41.M. Ravanelli, P. Brakel, M. Omologo, and Y. Bengio. Light gated recurrent units for speech recog- nition. IEEE Transactions on Emerging Topics in Computational Intelligence, 2(2):92–102, 2018.
  42. 42.M. Ravanelli, T. Parcollet, P. Plantinga, A. Rouhe, S. Cornell, L. Lugosch, C. Subakan, N. Dawalatabad, A. Heba, J. Zhong, et al. SpeechBrain: A general-purpose speech toolkit. CoRR, abs/2106.04624, 2021.
  43. 43.Y. Ren, C. Hu, X. Tan, T. Qin, S. Zhao, Z. Zhao, and T.-Y. Liu. Fastspeech 2: Fast and high- quality end-to-end text to speech. In Proceedings of the International Conference on Learning Representations (ICLR), 2021.
  44. 44.A. Rouhe, M. Ravanelli, T. Parcollet, and P. Plantinga. A SpeechBrain for Everything: State of the PyTorch Ecosystem for Speech Technologies. Interspeech Tutorial Presentation, September 2022.
  45. 45.J. Salazar, D. Liang, T. Q. Nguyen, and K. Kirchhoff. Masked language model scoring. CoRR, abs/1910.14659, 2019.
  46. 46.R. T. Schirrmeister, J. T. Springenberg, L. D. J. Fiederer, M. Glasstetter, K. Eggensperger, M. Tangermann, F. Hutter, W. Burgard, and T. Ball. Deep learning with convolutional neu- ral networks for eeg decoding and visualization. Human Brain Mapping, aug 2017a.
  47. 47.R. T. Schirrmeister, J. T. Springenberg, L. D. J. Fiederer, M. Glasstetter, K. Eggensperger, M. Tangermann, F. Hutter, W. Burgard, and T. Ball. Deep learning with convolutional neu- ral networks for EEG decoding and visualization. Human Brain Mapping, 38(11):5391–5420, Aug. 2017b.
  48. 48.J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. Skerry-Ryan, R. A. Saurous, Y. Agiomyrgiannakis, and Y. Wu. Natural TTS Synthesis by Conditioning WaveNet on Mel Spectrogram Predictions. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2017.
  49. 49.Y. Song, Q. Zheng, B. Liu, and X. Gao. EEG conformer: Convolutional transformer for EEG decoding and visualization. IEEE Transactions on Neural Systems and Rehabilitation Engineering, 31:710–719, 2023.
  50. 50.H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozi`ere, N. Goyal, E. Hambro, F. Azhar, A. Rodriguez, A. Joulin, E. Grave, and G. Lample. Llama: Open and efficient foundation language models. CoRR, abs/2302.13971, 2023.
  51. 51.A. D. Tur, A. Moumen, and M. Ravanelli. Progres: Prompted generative rescoring on asr n-best. In Proceedings of the IEEE Spoken Language Technology Workshop (SLT), 2024.
  52. 52.Y. Wang, M. Ravanelli, and A. Yacoubi. Speech Emotion Diarization: Which Emotion Appears When? In Proceedings of the IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), 2023.
  53. 53.S. Watanabe, T. Hori, S. Karita, T. Hayashi, J. Nishitoba, Y. Unno, N. Enrique Yalta Soplin, J. Heymann, M. Wiesner, N. Chen, A. Renduchintala, and T. Ochiai. ESPnet: End-to-end speech processing toolkit. In Proceedings of Interspeech, 2018.
  54. 54.S. wen Yang, P.-H. Chi, Y.-S. Chuang, C.-I. J. Lai, K. Lakhotia, Y. Y. Lin, A. T. Liu, J. Shi, X. Chang, G.-T. Lin, T.-H. Huang, W.-C. Tseng, K. tik Lee, D.-R. Liu, Z. Huang, S. Dong, S.-W. Li, S. Watanabe, A. Mohamed, and H. yi Lee. Superb: Speech processing universal performance benchmark. In Proceedings of Interspeech, 2021.
  55. 55.R. Whetten, T. Parcollet, M. Dinarelli, and Y. Est`eve. Open Implementation and Study of BEST- RQ for Speech Processing. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2024.
  56. 56.S. Zaiem, R. Algayres, T. Parcollet, E. Slim, and M. Ravanelli. Fine-tuning strategies for faster infer- ence using speech self-supervised models: A comparative study. In In Proceesings of the IEEE In- ternational Conference on Acoustics, Speech, and Signal Processing Workshops (ICASSP), 2023a.
  57. 57.S. Zaiem, Y. Kemiche, T. Parcollet, S. Essid, and M. Ravanelli. Speech Self-Supervised Represen- tation Benchmarking: Are We Doing it Right? In Proceedings of Interspeech, 2023b.

Citation

MLA
Ravanelli, M., et al. “Open-Source Conversational AI with SpeechBrain 1.0”. Journal of Machine Learning Research, vol. 25, no. 333, 2024, pp. 1–1, https://www.jmlr.org/papers/v25/24-0991.html.
APA
Ravanelli, M., Parcollet, T., Moumen, A., Langen, S. de ., Subakan, C., Plantinga, P., Wang, Y., Mousavi, P., Libera, L. D., Ploujnikov, A., Paissan, F., Borra, D., Zaiem, S., Zhao, Z., Zhang, S., Karakasidis, G., Yeh, S.-L., Champion, P., Rouhe, A., … Estève, Y. (2024). Open-Source Conversational AI with SpeechBrain 1.0. Journal of Machine Learning Research, 25(333), 1–11. https://www.jmlr.org/papers/v25/24-0991.html
Chicago
Ravanelli, M., T. Parcollet, A. Moumen, et al. 2024. “Open-Source Conversational AI with SpeechBrain 1.0”. Journal of Machine Learning Research 25 (333): 1–11. https://www.jmlr.org/papers/v25/24-0991.html.
Harvard
Ravanelli, M. et al. (2024) “Open-Source Conversational AI with SpeechBrain 1.0”, Journal of Machine Learning Research, 25(333), pp. 1–11. Available at: https://www.jmlr.org/papers/v25/24-0991.html.
Vancouver
1. Ravanelli M, Parcollet T, Moumen A, et al (2024) Open-Source Conversational AI with SpeechBrain 1.0. Journal of Machine Learning Research 25:1–11

BibTeX

@article{JMLR:v25:24-0991,
  author  = {Mirco Ravanelli and Titouan Parcollet and Adel Moumen and Sylvain de Langen and Cem Subakan and Peter Plantinga and Yingzhi Wang and Pooneh Mousavi and Luca Della Libera and Artem Ploujnikov and Francesco Paissan and Davide Borra and Salah Zaiem and Zeyu Zhao and Shucong Zhang and Georgios Karakasidis and Sung-Lin Yeh and Pierre Champion and Aku Rouhe and Rudolf Braun and Florian Mai and Juan Zuluaga-Gomez and Seyed Mahed Mousavi and Andreas Nautsch and Ha Nguyen and Xuechen Liu and Sangeet Sagar and Jarod Duret and Salima Mdhaffar and Ga{{\"e}}lle Laperri{{\`e}}re and Mickael Rouvier and Renato De Mori and Yannick Est{{\`e}}ve},
  title   = {Open-Source Conversational AI with SpeechBrain 1.0},
  journal = {Journal of Machine Learning Research},
  year    = {2024},
  volume  = {25},
  number  = {333},
  pages   = {1--11},
  url     = {http://jmlr.org/papers/v25/24-0991.html}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/