SoundStream: An End-to-End Neural Audio Codec
Neil ZeghidourAlejandro LuebsAhmed OmranJan SkoglundMarco Tagliasacchi
Proposes SoundStream, a real-time neural audio codec that pairs a convolutional autoencoder with residual vector quantization to achieve scalable, low-latency compression for general audio, outperforming traditional codecs at a fraction of their bitrate on mobile CPUs.
Real-time digital communication platforms and media streaming services face increasing demand for high-quality audio delivered over bandwidth-constrained and fluctuating networks. Traditional audio codecs (such as Opus and Enhanced Voice Services, or EVS) either rely on rigid signal-processing models tailored exclusively to speech at low bitrates or introduce noticeable distortion on diverse audio types like music when network capacity drops. The article presents SoundStream, a machine-learning-based audio compression system (neural codec) designed to deliver high-quality, general-purpose audio compression across variable bitrates while operating in real time with low latency.
The authors develop an end-to-end architecture comprising a fully convolutional encoder, a multi-stage residual vector quantizer, and a convolutional decoder. The system is trained jointly using a combination of adversarial and spectral reconstruction losses, alongside a novel "quantizer dropout" training technique that randomly varies the number of active quantizer stages. The model's performance was evaluated against industry standards (Opus, EVS, and Lyra) across clean speech, noisy speech, reverberant speech, and music sampled at 24 kHz, using both crowdsourced subjective listening tests and computational quality metrics.
The evaluation yielded several key findings. First, SoundStream at 3 kilobits per second (kbps) significantly outperformed Opus at 6 kbps and EVS at 5.9 kbps in perceptual listening tests; standard codecs required 3.2 to 4 times more bandwidth (9.6 kbps for EVS and 12 kbps for Opus) to match SoundStream’s 3 kbps audio quality. Second, unlike speech-only neural codecs, SoundStream successfully encoded music at 3 kbps with quality exceeding Opus at 12 kbps. Third, the quantizer dropout method allowed a single model to support dynamic bitrates between 3 kbps and 18 kbps with virtually no quality penalty compared to models trained specifically for a single bitrate. Fourth, the architecture achieved a low architectural latency of 13.3 milliseconds and executed over twice as fast as real time on a single smartphone CPU thread. Finally, integrating background noise suppression directly into the compression bottleneck delivered clean audio without increasing system latency or requiring a separate enhancement module.
These findings demonstrate that end-to-end neural audio codecs can replace traditional multi-stage pipelines and specialized codecs, substantially reducing network bandwidth costs and infrastructure complexity without degrading user experience. The ability to deploy a single lightweight model that dynamically adapts to network fluctuations and performs simultaneous noise reduction lowers memory and computational overhead for edge devices like mobile phones.
Organizations managing real-time communication platforms or audio streaming services should evaluate SoundStream as a next-generation replacement for legacy codecs in constrained network environments. Technical teams should conduct real-world pilot deployments to assess performance under live network conditions, such as packet loss and jitter. System architects can also take advantage of the asymmetric capacity finding—using a smaller encoder and larger decoder—to optimize battery and processing consumption on resource-constrained client devices.
The article's conclusions are supported by rigorous subjective and objective evaluations across multiple audio datasets. However, stakeholders should note that testing was limited to 24 kHz single-channel (mono) audio and focused on English-language speech corpora and specific music datasets. Additional validation is advised before deploying the architecture in multi-channel (spatial or stereo) setups, higher sampling rates (such as 48 kHz full-band audio), or unconstrained live acoustic environments with diverse languages and atypical background noise.
- Paper: Neural Discrete Representation Learning, Aäron van den Oord et al. (2017). This foundational paper introduces vector quantized variational autoencoders (VQ-VAE), establishing the discrete latent representation learning framework that SoundStream extends into residual vector quantization.
- Paper: HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis, Jungil Kong et al. (2020). It introduces the multi-period and multi-scale adversarial discriminator objectives that SoundStream directly adapts to synthesize high-fidelity raw audio waveforms from quantized embeddings.
- Paper: End-to-end Optimized Image Compression, Johannes Ballé et al. (2016). It establishes the foundational end-to-end optimization paradigm for learned data compression using non-linear transforms and rate-distortion objectives.
- Paper: Qwen3-TTS Technical Report, Hangrui Hu et al. (2026). This technical report builds on neural audio codec tokenizers like SoundStream to extract low-latency, discrete speech representations for large-scale generative text-to-speech models.
