SoundStream: An End-to-End Neural Audio Codec

Neil ZeghidourAlejandro LuebsAhmed OmranJan SkoglundMarco Tagliasacchi

article2021IEEE/ACM Transactions on Audio Speech and Language Processing1,496 citations

Proposes SoundStream, a real-time neural audio codec that pairs a convolutional autoencoder with residual vector quantization to achieve scalable, low-latency compression for general audio, outperforming traditional codecs at a fraction of their bitrate on mobile CPUs.

Listen

Real-time digital communication platforms and media streaming services face increasing demand for high-quality audio delivered over bandwidth-constrained and fluctuating networks. Traditional audio codecs (such as Opus and Enhanced Voice Services, or EVS) either rely on rigid signal-processing models tailored exclusively to speech at low bitrates or introduce noticeable distortion on diverse audio types like music when network capacity drops. The article presents SoundStream, a machine-learning-based audio compression system (neural codec) designed to deliver high-quality, general-purpose audio compression across variable bitrates while operating in real time with low latency.

The authors develop an end-to-end architecture comprising a fully convolutional encoder, a multi-stage residual vector quantizer, and a convolutional decoder. The system is trained jointly using a combination of adversarial and spectral reconstruction losses, alongside a novel "quantizer dropout" training technique that randomly varies the number of active quantizer stages. The model's performance was evaluated against industry standards (Opus, EVS, and Lyra) across clean speech, noisy speech, reverberant speech, and music sampled at 24 kHz, using both crowdsourced subjective listening tests and computational quality metrics.

The evaluation yielded several key findings. First, SoundStream at 3 kilobits per second (kbps) significantly outperformed Opus at 6 kbps and EVS at 5.9 kbps in perceptual listening tests; standard codecs required 3.2 to 4 times more bandwidth (9.6 kbps for EVS and 12 kbps for Opus) to match SoundStream’s 3 kbps audio quality. Second, unlike speech-only neural codecs, SoundStream successfully encoded music at 3 kbps with quality exceeding Opus at 12 kbps. Third, the quantizer dropout method allowed a single model to support dynamic bitrates between 3 kbps and 18 kbps with virtually no quality penalty compared to models trained specifically for a single bitrate. Fourth, the architecture achieved a low architectural latency of 13.3 milliseconds and executed over twice as fast as real time on a single smartphone CPU thread. Finally, integrating background noise suppression directly into the compression bottleneck delivered clean audio without increasing system latency or requiring a separate enhancement module.

These findings demonstrate that end-to-end neural audio codecs can replace traditional multi-stage pipelines and specialized codecs, substantially reducing network bandwidth costs and infrastructure complexity without degrading user experience. The ability to deploy a single lightweight model that dynamically adapts to network fluctuations and performs simultaneous noise reduction lowers memory and computational overhead for edge devices like mobile phones.

Organizations managing real-time communication platforms or audio streaming services should evaluate SoundStream as a next-generation replacement for legacy codecs in constrained network environments. Technical teams should conduct real-world pilot deployments to assess performance under live network conditions, such as packet loss and jitter. System architects can also take advantage of the asymmetric capacity finding—using a smaller encoder and larger decoder—to optimize battery and processing consumption on resource-constrained client devices.

The article's conclusions are supported by rigorous subjective and objective evaluations across multiple audio datasets. However, stakeholders should note that testing was limited to 24 kHz single-channel (mono) audio and focused on English-language speech corpora and specific music datasets. Additional validation is advised before deploying the architecture in multi-channel (spatial or stereo) setups, higher sampling rates (such as 48 kHz full-band audio), or unconstrained live acoustic environments with diverse languages and atypical background noise.

arXiv: 2107.03312
  • Paper: Neural Discrete Representation Learning, Aäron van den Oord et al. (2017). This foundational paper introduces vector quantized variational autoencoders (VQ-VAE), establishing the discrete latent representation learning framework that SoundStream extends into residual vector quantization.
  • Paper: HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis, Jungil Kong et al. (2020). It introduces the multi-period and multi-scale adversarial discriminator objectives that SoundStream directly adapts to synthesize high-fidelity raw audio waveforms from quantized embeddings.
  • Paper: End-to-end Optimized Image Compression, Johannes Ballé et al. (2016). It establishes the foundational end-to-end optimization paradigm for learned data compression using non-linear transforms and rate-distortion objectives.
  • Paper: Qwen3-TTS Technical Report, Hangrui Hu et al. (2026). This technical report builds on neural audio codec tokenizers like SoundStream to extract low-latency, discrete speech representations for large-scale generative text-to-speech models.
Cover for SoundStream: An End-to-End Neural Audio Codec

Abstract

We present SoundStream, a novel neural audio codec that can efficiently compress speech, music and general audio at bitrates normally targeted by speech-tailored codecs. SoundStream relies on a model architecture composed by a fully convolutional encoder/decoder network and a residual vector quantizer, which are trained jointly end-to-end. Training leverages recent advances in text-to-speech and speech enhancement, which combine adversarial and reconstruction losses to allow the generation of high-quality audio content from quantized embeddings. By training with structured dropout applied to quantizer layers, a single model can operate across variable bitrates from 3kbps to 18kbps, with a negligible quality loss when compared with models trained at fixed bitrates. In addition, the model is amenable to a low latency implementation, which supports streamable inference and runs in real time on a smartphone CPU. In subjective evaluations using audio at 24kHz sampling rate, SoundStream at 3kbps outperforms Opus at 12kbps and approaches EVS at 9.6kbps. Moreover, we are able to perform joint compression and enhancement either at the encoder or at the decoder side with no additional latency, which we demonstrate through background noise suppression for speech.

Table of Contents

  • I Introduction
  • II Related work
  • III Model
  • III-A Encoder architecture
  • III-B Decoder architecture
  • III-C Residual Vector Quantizer:
  • III-D Discriminator architecture
  • III-E Training objective
  • III-F Joint compression and enhancement
  • IV Evaluation setup
  • IV-A Datasets
  • IV-B Evaluation metrics
  • IV-C Baselines
  • V Results
  • V-A Comparison with other codecs
  • V-B Objective quality metrics
  • V-C Bitrate scalability
  • V-D Ablation studies
  • V-E Joint compression and enhancement
  • V-F Joint vs. disjoint compression and enhancement
  • VI Conclusions
  • References

Knowls

  1. Knowl 1 — SoundStream End-to-End Neural Audio Codec Architecture

    model/method

    SoundStream is an end-to-end neural audio codec designed for compression of speech, music, and general audio at 24 kHz sampling rate across bitrates from 3 kbps to 18 kbps. It maps a single-channel audio waveform x∈RTx \in \mathbb{R}^T to quantized discrete representations and decodes a lossy reconstruction x^∈RT\hat{x} \in \mathbb{R}^T in real time with low latency.

    The system consists of three causal building blocks:

    1. Causal Convolutional Encoder: Maps time-domain waveform xx to latent embedding sequence enc(x)∈RS×D\text{enc}(x) \in \mathbb{R}^{S \times D}, where S=T/MS = T / M and MM is the total downsampling ratio. It begins with a 1D convolution with CencC_{\text{enc}} channels and kernel size 7, followed by Benc=4B_{\text{enc}} = 4 convolutional blocks with striding sequence (2,4,5,8)(2, 4, 5, 8), yielding M=320M = 320 samples per frame (a frame rate of S/T=75 HzS/T = 75\text{ Hz} at 24 kHz, or 13.3 ms13.3\text{ ms} per embedding). Each block contains three residual units with dilated 1D convolutions (kernel size 7, dilation rates 1, 3, and 9) followed by a strided downsampling convolution where channel count is doubled. A final 1D convolution with kernel size 3 and stride 1 outputs DD-dimensional embeddings.
    2. Residual Vector Quantizer (RVQ): Discretizes the sequence of continuous DD-dimensional embeddings using NqN_q cascaded vector quantizers.
    3. Causal Convolutional Decoder: Mirrors the encoder structure to reconstruct the waveform x^\hat{x} from quantized embeddings y^∈RS×D\hat{y} \in \mathbb{R}^{S \times D}. It consists of a 1D convolution with kernel size 7 and 16Cdec16 C_{\text{dec}} channels, followed by Bdec=4B_{\text{dec}} = 4 transposed convolution blocks with upsampling strides (8,5,4,2)(8, 5, 4, 2) where channel counts halve at each block. Each block contains the same three residual units (dilation rates 1, 3, and 9). A final 1D convolution with 1 output channel and kernel size 7 projects back to the waveform domain.

    All convolutions in the encoder and decoder are causal (padding is applied only to past samples during training and offline inference, and no padding is used during streaming inference), enabling streamable inference with architectural latency determined solely by the resampling stride MM.

  2. Knowl 2 — Residual Vector Quantization Algorithm and Codebook Training

    algorithm

    Residual Vector Quantization (RVQ) compresses continuous DD-dimensional latent embeddings y=enc(x)∈RS×Dy = \text{enc}(x) \in \mathbb{R}^{S \times D} produced at frame rate SS by cascading NqN_q discrete vector quantizer layers Q1,…,QNqQ_1, \dots, Q_{N_q}. Each quantizer QiQ_i maintains a codebook of NN vectors in RD\mathbb{R}^D, allocating ri=log⁡2Nr_i = \log_2 N bits per frame and yielding a target bitrate of R=S⋅Nqlog⁡2NR = S \cdot N_q \log_2 N bits per second.

    Input: Continuous encoder output y∈RS×Dy \in \mathbb{R}^{S \times D}, vector quantizers QiQ_i for i=1,…,Nqi = 1, \dots, N_q
    Output: Quantized representation y^∈RS×D\hat{y} \in \mathbb{R}^{S \times D}
    y^←0.0\hat{y} \leftarrow 0.0
    residual←y\text{residual} \leftarrow y
    for i=1i = 1 to NqN_q do
        y^←y^+Qi(residual)\hat{y} \leftarrow \hat{y} + Q_i(\text{residual})
        residual←residual−Qi(residual)\text{residual} \leftarrow \text{residual} - Q_i(\text{residual})
    return y^\hat{y}

    Codebook optimization and vector usage are managed through three complementary techniques:

    1. Exponential Moving Average (EMA) Updates: Codebook centroids are updated across training batches via exponential moving averages.
    2. K-Means Initialization: Centroids for all codebooks are initialized using the k-means clustering algorithm run on the first training batch.
    3. Dead Code Replacement: An exponential moving average of assignments to each codebook vector is tracked with a decay factor of 0.99. If this assignment count drops below 2, the underused codebook vector is replaced with an input frame randomly sampled from the current batch.
  3. Knowl 3 — Structured Quantizer Dropout for Variable Bitrate Scalability

    model/method

    Bitrate scalability in the SoundStream neural codec is achieved using structured quantizer dropout during training. Instead of training separate models for different target bitrates, a single model serves a continuous range of bitrates corresponding to nq∈{1,…,Nq}n_q \in \{1, \dots, N_q\} active quantizer layers.

    During training, for each input audio example, the number of active quantizers nqn_q is sampled uniformly at random from {1,…,Nq}\{1, \dots, N_q\}. The quantized latent embedding is computed using only the first nqn_q residual vector quantizer stages:

    y^=∑i=1nqQi(residuali−1)\hat{y} = \sum_{i=1}^{n_q} Q_i(\text{residual}_{i-1})

    The decoder is optimized to reconstruct the audio from representations quantized at any depth nqn_q. Because the outputs of all RVQ stages are additively combined in RS×D\mathbb{R}^{S \times D}, the embedding dimension and decoder architecture remain identical across all bitrates. During inference, the transmitter selects nqn_q according to the desired transmission bitrate R=S⋅nqlog⁡2NR = S \cdot n_q \log_2 N without requiring retraining or architecture modification.

  4. Knowl 4 — SoundStream Multi-Objective Training Loss

    equation

    SoundStream is trained end-to-end using a weighted combination of adversarial, feature matching, and multi-scale spectral reconstruction losses. Let G(x)=dec(Q(enc(x)))G(x) = \text{dec}(Q(\text{enc}(x))) denote the generator output for audio waveform xx, and let k∈{0,…,K}k \in \{0, \dots, K\} index the discriminators (where k=0k=0 is an STFT-based discriminator and k∈{1,…,K}k \in \{1, \dots, K\} represent K=3K=3 scales of a waveform-based discriminator). TkT_k is the number of time logits in discriminator kk, and Dk,tD_{k,t} denotes its output at time step tt.

    The discriminators minimize the hinge loss:

    LD=Ex[1K+1∑k=0K1Tk∑tmax⁡(0,1−Dk,t(x))+1K+1∑k=0K1Tk∑tmax⁡(0,1+Dk,t(G(x)))]\mathcal{L}_D = \mathbb{E}_x \left[ \frac{1}{K+1} \sum_{k=0}^K \frac{1}{T_k} \sum_t \max(0, 1 - D_{k,t}(x)) + \frac{1}{K+1} \sum_{k=0}^K \frac{1}{T_k} \sum_t \max(0, 1 + D_{k,t}(G(x))) \right]

    The generator GG minimizes:

    LG=λadvLGadv+λfeatLGfeat+λrecLGrec\mathcal{L}_G = \lambda_{\text{adv}} \mathcal{L}_G^{\text{adv}} + \lambda_{\text{feat}} \mathcal{L}_G^{\text{feat}} + \lambda_{\text{rec}} \mathcal{L}_G^{\text{rec}}

    with hyperparameters λadv=1\lambda_{\text{adv}} = 1, λfeat=100\lambda_{\text{feat}} = 100, and λrec=1\lambda_{\text{rec}} = 1, where:

    LGadv=Ex[1K+1∑k=0K1Tk∑tmax⁡(0,1−Dk,t(G(x)))]\mathcal{L}_G^{\text{adv}} = \mathbb{E}_x \left[ \frac{1}{K+1} \sum_{k=0}^K \frac{1}{T_k} \sum_t \max(0, 1 - D_{k,t}(G(x))) \right]

    LGfeat=Ex[1(K+1)L∑k=0K∑l=1L1Tk,l∑t∣Dk,t(l)(x)−Dk,t(l)(G(x))∣]\mathcal{L}_G^{\text{feat}} = \mathbb{E}_x \left[ \frac{1}{(K+1)L} \sum_{k=0}^K \sum_{l=1}^L \frac{1}{T_{k,l}} \sum_t \left| D_{k,t}^{(l)}(x) - D_{k,t}^{(l)}(G(x)) \right| \right]

    where LL is the number of internal layers and Dk,t(l)D_{k,t}^{(l)} is the layer ll activation at time tt with length Tk,lT_{k,l}.

    The multi-scale spectral reconstruction loss is:

    LGrec=∑s∈{26,27,28,29,210,211}∑t(∥Sts(x)−Sts(G(x))∥1+s2∥log⁡Sts(x)−log⁡Sts(G(x))∥2)\mathcal{L}_G^{\text{rec}} = \sum_{s \in \{2^6, 2^7, 2^8, 2^9, 2^{10}, 2^{11}\}} \sum_t \left( \left\| S_t^s(x) - S_t^s(G(x)) \right\|_1 + \sqrt{\frac{s}{2}} \left\| \log S_t^s(x) - \log S_t^s(G(x)) \right\|_2 \right)

    where Sts(x)S_t^s(x) is the tt-th frame of a 64-bin mel-spectrogram computed with STFT window length ss and hop size s/4s/4.

  5. Knowl 5 — Dual Discriminator Architecture: Waveform and STFT Discriminators

    model/method

    To guide generative training, SoundStream uses two complementary discriminator architectures comprising four separate discriminator networks (K+1=4K+1=4):

    1. Wave-Based Multi-Resolution Discriminator (K=3K=3): Three structurally identical convolutional networks process the audio waveform at original (1×1\times), 2×2\times downsampled, and 4×4\times downsampled resolutions. Each network consists of an initial 1D convolution followed by four grouped convolutions (group size 4, stride 4, channel multiplier 4 up to a cap of 1024 channels) and two concluding 1D convolution layers that output time-domain logits.
    2. STFT-Based Time-Frequency Discriminator (k=0k=0): Operates on the complex-valued Short-Time Fourier Transform (real and imaginary parts) of the waveform computed with window length W=1024W = 1024 samples and hop length H=256H = 256 samples (F=W/2=512F = W/2 = 512 frequency bins). The network consists of an initial 2D convolution (kernel 7×77 \times 7, 32 channels), followed by 6 residual blocks with 3×33 \times 3 convolutions and downsampling convolutions (3×43 \times 4 or 4×44 \times 4) that alternate between strides (st,sf)=(1,2)(s_t, s_f) = (1, 2) and (2,2)(2, 2) in time and frequency. The resulting representation of shape TH⋅23×F26\frac{T}{H \cdot 2^3} \times \frac{F}{2^6} is aggregated across downsampled frequency bins by a 1×(F/26)1 \times (F/2^6) 2D convolution to yield a 1D sequence of logits.
  6. Knowl 6 — Joint Audio Compression and Controllable Noise Suppression via FiLM Conditioning

    model/method

    SoundStream integrates audio compression and background noise suppression into a single model without increasing architectural latency or buffering requirements. Controllable denoising is achieved by applying Feature-wise Linear Modulation (FiLM) layers at the bottleneck (either on the encoder embeddings before quantization or on the decoder embeddings after quantization).

    The FiLM transformation modulates intermediate activation an,ca_{n,c} (nn-th sample, cc-th channel) as:

    a~n,c=γn,can,c+βn,c\tilde{a}_{n,c} = \gamma_{n,c} a_{n,c} + \beta_{n,c}

    where γn,c\gamma_{n,c} and βn,c\beta_{n,c} are produced by a linear layer conditioned on a 2D one-hot vector indicating whether denoising is enabled or disabled.

    The model is trained on data tuples of (inputs, targets, denoise):

    • When denoise = false, targets = inputs.
    • When denoise = true on noisy speech, targets contains only the clean speech component.
    • When inputs consists of clean speech or music, targets = inputs regardless of the denoise flag, preventing distortion of clean audio.

    Applying conditioning at the encoder before quantization reduces the cross-entropy of codebook assignments, creating a representation that yields larger potential bitrate savings (between 7% and 20%) when entropy coding is applied.

  7. Knowl 7 — Subjective Audio Quality of SoundStream Across Bitrates and Content Types

    empirical result

    In MUSHRA-style crowdsourced listening tests on 24 kHz audio covering clean speech, noisy speech, reverberant speech, and music:

    • Low Bitrates (3 kbps): SoundStream operating at 3 kbps significantly outperforms Opus at 6 kbps and 12 kbps, EVS at 5.9 kbps, and the Lyra neural codec at 3 kbps. It approaches the quality of EVS at 9.6 kbps, matching or exceeding standard codecs that require 3.2×3.2\times to 4×4\times higher bitrates (e.g., Opus at 12 kbps).
    • Medium Bitrates (6 kbps): SoundStream at 6 kbps matches the subjective quality of EVS at 13.2 kbps and Opus at 16 kbps, achieving a 2.2×2.2\times to 2.6×2.6\times bitrate reduction.
    • High Bitrates (12 kbps): SoundStream at 12 kbps matches EVS at 16.4 kbps and Opus at 20 kbps (1.3×1.3\times to 1.6×1.6\times bitrate reduction).
    • Content Generalization: SoundStream maintains consistent quality across clean and noisy/reverberant speech, and encodes music at 3 kbps with subjective quality substantially superior to Opus at 12 kbps and EVS at 5.9 kbps.
  8. Knowl 8 — Rate-Quality Advantage of a Learned Convolutional Encoder over Fixed Spectral Features

    empirical result

    Replacing the learnable causal convolutional encoder in SoundStream with a fixed mel-filterbank representation (while keeping the trainable residual vector quantizer and decoder) causes a severe degradation in audio quality: at 6 kbps, the objective quality score measured by ViSQOL drops from 3.96 (learned encoder) to 3.33 (fixed mel-filterbank).

    A learned-encoder SoundStream model operating at half the bitrate (3 kbps) achieves a ViSQOL score of 3.76, significantly outperforming the fixed mel-filterbank model operating at 6 kbps (3.33). This demonstrates that end-to-end optimization of the encoder is the critical factor enabling high coding efficiency at low bitrates.

  9. Knowl 9 — Computational Efficiency, Quantizer Depth, and Latency Trade-Offs

    data/table

    Ablation experiments evaluated on 24 kHz audio at 6 kbps examine the effects of encoder/decoder capacity, RVQ depth/codebook size, and striding latency on ViSQOL quality and Real-Time Factor (RTF, audio duration divided by processing time) profiled on a single CPU thread of a Google Pixel 4 smartphone.

    CencC_{\text{enc}} CdecC_{\text{dec}} #Params RTF (enc) RTF (dec) ViSQOL
    32 32 8.4 M 2.4×2.4\times 2.3×2.3\times 4.01±0.034.01 \pm 0.03
    16 16 2.4 M 7.5×7.5\times 7.1×7.1\times 3.98±0.033.98 \pm 0.03
    16 32 5.5 M 7.5×7.5\times 2.3×2.3\times 4.02±0.034.02 \pm 0.03
    8 32 4.8 M 18.6×18.6\times 2.3×2.3\times 3.99±0.033.99 \pm 0.03
    32 16 5.3 M 2.4×2.4\times 7.1×7.1\times 3.97±0.033.97 \pm 0.03
    32 8 4.4 M 2.4×2.4\times 17.1×17.1\times 3.90±0.033.90 \pm 0.03
    Number of quantizers NqN_q 8 16 80
    Codebook size NN 1024 32 2
    ViSQOL 4.01±0.034.01 \pm 0.03 3.98±0.033.98 \pm 0.03 3.92±0.033.92 \pm 0.03
    Strides Latency NqN_q RTF (enc) RTF (dec) ViSQOL
    (1, 4, 5, 8) 7.5 ms 4 1.6×1.6\times 1.5×1.5\times 4.01±0.024.01 \pm 0.02
    (2, 4, 5, 8) 13.3 ms 8 2.4×2.4\times 2.3×2.3\times 4.01±0.034.01 \pm 0.03
    (4, 4, 5, 8) 26.6 ms 16 4.1×4.1\times 4.0×4.0\times 4.01±0.034.01 \pm 0.03

    Key takeaways:

    1. Asymmetric capacity: Reducing encoder capacity (Cenc=8,Cdec=32C_{\text{enc}}=8, C_{\text{dec}}=32) accelerates encoder execution to 18.6×18.6\times real-time with negligible quality loss (3.993.99 vs 4.014.01), whereas reducing decoder capacity causes a significant drop (3.903.90).
    2. Quantizer depth: An RVQ cascade of 80 1-bit quantizers (N=2N=2) trains stably and achieves 3.923.92 ViSQOL, demonstrating that RVQ scales gracefully without optimization failure.
    3. Latency: Audio quality is invariant across frame latencies of 7.5 ms, 13.3 ms, and 26.6 ms when adjusting NqN_q to maintain a 6 kbps budget, while longer frame strides yield higher real-time factors.
  10. Knowl 10 — Joint versus Disjoint Audio Compression and Speech Enhancement

    data/table

    The performance of a single SoundStream model performing joint compression (at 6 kbps) and speech enhancement is compared against two disjoint pipelines combining SoundStream (denoising disabled) and a dedicated SEANet enhancement model. Evaluations were conducted on 1000 noisy speech clips from the VCTK corpus across input Signal-to-Noise Ratios (SNRs).

    Input SNR Joint SoundStream SoundStream →\rightarrow SEANet SEANet →\rightarrow SoundStream
    0 dB 2.93±0.022.93 \pm 0.02 3.02±0.033.02 \pm 0.03 3.05±0.023.05 \pm 0.02
    5 dB 3.18±0.023.18 \pm 0.02 3.30±0.023.30 \pm 0.02 3.31±0.023.31 \pm 0.02
    10 dB 3.42±0.023.42 \pm 0.02 3.51±0.023.51 \pm 0.02 3.50±0.023.50 \pm 0.02
    15 dB 3.58±0.023.58 \pm 0.02 3.64±0.023.64 \pm 0.02 3.63±0.023.63 \pm 0.02

    The unified SoundStream model achieves enhancement quality nearly matching dedicated sequential pipelines (with the gap narrowing as input SNR increases), while cutting computational cost in half and eliminating the extra architectural latency introduced by stacking separate processing models.

Coverage note — None omitted; all primary contributions, including architecture definitions, quantization algorithms, training loss formulations, conditioning mechanisms, and empirical/ablation evaluations, are represented.

References

  1. 1.Y. Li, M. Tagliasacchi, O. Rybakov, V. Ungureanu, and D. Roblek, “Real-time speech frequency bandwidth extension,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2021, pp. 691–695.
  2. 2.A. Biswas and D. Jia, “ societies codec enhancement with generative adversarial networks,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2020, pp. 356–360.
  3. 3.F. Stimberg, A. Narest, A. Bazzica, L. Kolmodin, P. Barrera González, O. Sharonova, H. Lundin, and T. C. Walters, “WaveNetEQ — Packet loss concealment with WaveRNN,” in 54th Asilomar Conference on Signals, Systems, and Computers, 2020, pp. 672–676.
  4. 4.A. van den Oord, S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A. Graves, N. Kalchbrenner, A. Senior, and K. Kavukcuoglu, “WaveNet: A generative model for raw audio,” arXiv:1609.03499, 2016.
  5. 5.W. B. Kleijn, F. S. Lim, A. Luebs, J. Skoglund, F. Stimberg, Q. Wang, and T. C. Walters, “WaveNet based low rate speech coding,” in IEEE international conference on acoustics, speech and signal processing (ICASSP), 2018, pp. 676–680.
  6. 6.C. Gârbacea, A. van den Oord, Y. Li, F. S. C. Lim, A. Luebs, O. Vinyals, and T. C. Walters, “ societies bit-rate speech coding with VQ-VAE and a WaveNet decoder,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2019, pp. 735–739.
  7. 7.J.-M. Valin and J. Skoglund, “ LPCNet: improving neural speech synthesis through linear prediction,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2019, pp. 5891–5895.
  8. 8.W. B. Kleijn, A. Storus, M. Chinen, T. Denton, F. S. C. Lim, A. Luebs, J. Skoglund, and H. Yeh, “Generative speech coding with predictive variance regularization,” arXiv:2102.09660, 2021.
  9. 9.J.-M. Valin, K. Vos, and T. B. Terriberry, “Definition of the Opus Audio Codec,” IETF RFC 6716, 2012, https://tools.ietf.org/html/rfc6716.
  10. 10.M. Dietz, M. Multrus, V. Eksler, V. Malenovsky, E. Norvell, H. Pobloth, L. Miao, Z. Wang, L. Laaksonen, A. Vasilache, Y. Kamamoto, K. Kikuiri, S. Ragot, J. Faure, H. Ehara, V. Rajendran, V. Atti, H. Sung, E. Oh, H. Yuan, and C. Zhu, “ societies of the EVS codec architecture,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2015, pp. 5698–5702.
  11. 11.S. Mehri, K. Kumar, I. Gulrajani, R. Kumar, S. Jain, J. Sotelo, A. Courville, and Y. Bengio, “ societies: An unconditional end-to-end neural audio generation model,” arXiv:1612.07837, 2017.
  12. 12.A. van den Oord, Y. Li, I. Babuschkin, K. Simonyan, O. Vinyals, K. Kavukcuoglu, G. van den Driessche, E. Lockhart, L. Cobo, F. Stimberg, N. Casagrande, D. Grewe, S. Noury, S. Dieleman, E. Elsen, N. Kalchbrenner, H. Zen, A. Graves, H. King, T. Walters, D. Belov, and D. Hassabis, “ societies: Fast high-fidelity speech synthesis,” in Proceedings of the 35th International Conference on Machine Learning, 2018, pp. 3918–3926.
  13. 13.N. Kalchbrenner, E. Elsen, K. Simonyan, S. Noury, N. Casagrande, E. Lockhart, F. Stimberg, A. van den Oord, S. Dieleman, and K. Kavukcuoglu, “ Efficient neural audio synthesis,” arXiv:1802.08435, 2018.
  14. 14.Z. Jin, A. Finkelstein, G. J. Mysore, and J. Lu, “ FFTNet: a real-time speaker-dependent neural vocoder,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2018, pp. 2251–2255.
  15. 15.K. Kumar, R. Kumar, T. de Boissiere, L. Gestin, W. Z. Teoh, J. Sotelo, A. de Brebisson, Y. Bengio, and A. Courville, “ MelGAN: Generative adversarial networks for conditional waveform synthesis,” in Advances in Neural Information Processing Systems, 2019.
  16. 16.J. Kong, J. Kim, and J. Bae, “ HiFi-GAN: Generative Adversarial Networks for efficient and high fidelity speech synthesis,” arXiv:2010.05646, 2020.
  17. 17.X. Feng, Y. Zhang, and J. Glass, “ societies feature denoising and dereverberation via deep autoencoders for noisy reverberant speech recognition,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2014, pp. 1759–1763.
  18. 18.S. Pascual, A. Bonafonte, and J. Serra, “ SEGAN: Speech enhancement generative adversarial network,” arXiv:1703.09452, 2017.
  19. 19.F. G. Germain, Q. Chen, and V. Koltun, “ societies denoising with deep feature losses,” arXiv:1806.10522, 2018.
  20. 20.D. Rethage, J. Pons, and X. Serra, “ A WaveNet for speech denoising,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2018, pp. 5069–5073.
  21. 21.C. Donahue, B. Li, and R. Prabhavalkar, “ Exploring speech enhancement with generative adversarial networks for robust speech recognition,” in 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2018, pp. 5024–5028.
  22. 22.T. Ishii, H. Komiyama, T. Shinozaki, Y. Horiuchi, and S. Kuroiwa, “ Reverberant speech recognition based on denoising autoencoder.” in Interspeech, 2013, pp. 3512–3516.
  23. 23.D. S. Williamson and D. Wang, “ Time-frequency masking in the complex domain for speech dereverberation and denoising,” IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 25, pp. 1492–1501, 2017.
  24. 24.T. Y. Lim, R. A. Yeh, Y. Xu, M. N. Do, and M. Hasegawa-Johnson, “ Time-frequency networks for audio super-resolution,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2018, pp. 646–650.
  25. 25.S. Lloyd, “ Least squares quantization in PCM,” IEEE transactions on information theory, vol. 28, pp. 129–137, 1982.
  26. 26.Y. Linde, A. Buzo, and R. Gray, “ An algorithm for vector quantizer design,” IEEE Transactions on Communications, vol. 28, pp. 84–95, 1980.
  27. 27.J. MacQueen, “ Some methods for classification and analysis of multivariate observations,” Proceedings of the Fifth Berkeley Symposium on Mathematical Statistics and Probability, pp. 281–297, 1967.
  28. 28.R. Gray, “ Vector quantization,” IEEE ASSP Magazine, vol. 1, pp. 4–29, 1984.
  29. 29.J. Makhoul, S. Roucos, and H. Gish, “ Vector quantization in speech coding,” Proceedings of the IEEE, vol. 73, pp. 1551–1588, 1985.
  30. 30.M. Schroeder and B. Atal, “ Code-excited linear prediction (CELP): High-quality speech at very low bit rates,” in IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP), 1985, pp. 937–940.
  31. 31.A. van den Oord, O. Vinyals, and K. Kavukcuoglu, “ Neural discrete representation learning,” arXiv:1711.00937, 2017.
  32. 32.A. Razavi, A. van den Oord, and O. Vinyals, “ Generating diverse high-fidelity images with VQ-VAE-2,” arXiv:1906.00446, 2019.
  33. 33.S. Dieleman, A. van den Oord, and K. Simonyan, “ The challenge of realistic music generation: Modelling raw audio at scale,” arXiv:1806.10474, 2018.
  34. 34.P. Dhariwal, H. Jun, C. Payne, J. W. Kim, A. Radford, and I. Sutskever, “ Jukebox: A generative model for music,” arXiv:2005.00341, 2020.
  35. 35.B.-H. Juang and A. Gray, “ Multiple stage vector quantization for speech coding,” in IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP), 1982, pp. 597–600.
  36. 36.A. Vasuki and P. Vanathi, “ A review of vector quantization techniques,” IEEE Potentials, vol. 25, pp. 39–47, 2006.
  37. 37.S. Morishima, H. Harashima, and Y. Katayama, “ Speech coding based on a multi-layer neural network,” in IEEE International Conference on Communications, Including Supercomm Technical Sessions, 1990, pp. 429–433.
  38. 38.S. Kankanahalli, “ End-to-end optimized speech coding with deep neural networks,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2018, pp. 2521–2525.
  39. 39.A. Polyak, Y. Adi, J. Copet, E. Kharitonov, K. Lakhotia, W.-N. Hsu, A. Mohamed, and E. Dupoux, “ Speech resynthesis from discrete disentangled self-supervised representations,” arXiv:2104.00355, 2021.
  40. 40.K. Zhen, J. Sung, M. S. Lee, S. Beack, and M. Kim, “ Cascaded cross-module residual learning towards lightweight end-to-end speech coding,” arXiv:1906.07769, 2019.
  41. 41.J. Casebeer, V. Vale, U. Isik, J.-M. Valin, R. Giri, and A. Krishnaswamy, “ Enhancing into the codec: Noise robust speech coding with vector-quantized autoencoders,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2021, pp. 711–715.
  42. 42.D.-A. Clevert, T. Unterthiner, and S. Hochreiter, “ Fast and accurate deep network learning by exponential linear units (elus),” arXiv preprint arXiv:1511.07289, 2015.
  43. 43.N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov, “ Dropout: a simple way to prevent neural networks from overfitting,” The journal of machine learning research, vol. 15, pp. 1929–1958, 2014.
  44. 44.A. Baevski, H. Zhou, A. Mohamed, and M. Auli, “ wav2vec 2.0: A framework for self-supervised learning of speech representations,” arXiv:2006.11477, 2020.
  45. 45.M. Tagliasacchi, Y. Li, K. Misiunas, and D. Roblek, “ SEANet: A multi-modal speech enhancement network,” in Interspeech, 2020, pp. 1126–1130.
  46. 46.Y. Blau and T. Michaeli, “ The perception-distortion tradeoff,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2018, pp. 6228–6237.
  47. 47.J. Engel, L. H. Hantrakul, C. Gu, and A. Roberts, “ DDSP: Differentiable digital signal processing,” arXiv:2001.04643, 2020.
  48. 48.A. A. Gritsenko, T. Salimans, R. van den Berg, J. Snoek, and N. Kalchbrenner, “ A spectral energy distance for parallel speech synthesis,” arXiv:2008.01160, 2020.
  49. 49.E. Perez, F. Strub, H. de Vries, V. Dumoulin, and A. Courville, “ FiLM: Visual reasoning with a general conditioning layer,” Proceedings of the AAAI Conference on Artificial Intelligence, vol. 32, pp. 3942–3951, 2018.
  50. 50.H. Zen, V. Dang, R. Clark, Y. Zhang, R. J. Weiss, Y. Jia, Z. Chen, and Y. Wu, “ LibriTTS: a corpus derived from LibriSpeech for text-to-speech,” arXiv:1904.02882, 2019.
  51. 51.E. Fonseca, J. Pons Puig, X. Favory, F. Font Corbera, D. Bogdanov, A. Ferraro, S. Oramas, A. Porter, and X. Serra, “ Freesound datasets: a platform for the creation of open audio datasets,” in Proceedings of the 18th ISMIR Conference, 2017, pp. 486–493.
  52. 52.E. Law, K. West, M. I. Mandel, M. Bay, and J. S. Downie, “ Evaluation of algorithms using games: The case of music tagging.” in ISMIR, 2009, pp. 387–392.
  53. 53.ITU-R, Recommendation BS.1534-1: Method for the subjective assessment of intermediate quality level of coding systems, International Telecommunications Union, 2001.
  54. 54.ITU, “ Perceptual evaluation of speech quality (PESQ): an objective method for end-to-end speech quality assessment of narrow-band telephone networks and speech codecs,” Int. Telecomm. Union, Geneva, Switzerland, ITU-T Rec. P.862, 2001.
  55. 55.——, “Perceptual objective listening quality assessment,” Int. Telecomm. Union, Geneva, Switzerland, ITU-T Rec. P.863, 2018.
  56. 56.A. Hines, J. Skoglund, A. Kokaram, and N. Harte, “ ViSQOL: The virtual speech quality objective listener,” in International Workshop on Acoustic Signal Enhancement (IWAENC), 2012, pp. 1–4.
  57. 57.M. Chinen, F. S. C. Lim, J. Skoglund, N. Gureev, F. O’Gorman, and A. Hines, “ ViSQOL v3: an open source production ready objective speech and audio metric,” in Twelfth International Conference on Quality of Multimedia Experience (QoMEX), 2020, pp. 1–6.
  58. 58.W3C, “ WebRTC 1.0: Real-time communication between browsers,” 2019, https://www.w3.org/TR/webrtc/.
  59. 59.C. Holmberg, S. Håkansson, and G. Eriksson, “ Web real-time communication use cases and requirements,” IETF RFC 7478, Mar. 2015, https://tools.ietf.org/html/rfc7478.
  60. 60.B. Bessette, R. Salami, R. Lefebvre, M. Jelinek, J. Rotola-Pukkila, J. Vainio, H. Mikkola, and K. Järvinen, “ The adaptive multirate wideband speech codec (AMR-WB),” IEEE Transactions on Speech and Audio Processing, vol. 10, pp. 620–636, 2002.
  61. 61.F. Mentzer, G. Toderici, M. Tschannen, and E. Agustsson, “ High-fidelity generative image compression,” arXiv:2006.09965, 2020.
  62. 62.J. Yamagishi, C. Veaux, and K. MacDonald, “ CSTR VCTK Corpus: English multi-speaker corpus for CSTR voice cloning toolkit (version 0.92),” 2019.

Citation

MLA
Zeghidour, N., et al. “SoundStream: An End-to-End Neural Audio Codec”. arXiv, 2021, http://arxiv.org/abs/2107.03312v1.
APA
Zeghidour, N., Luebs, A., Omran, A., Skoglund, J., & Tagliasacchi, M. (2021). SoundStream: An End-to-End Neural Audio Codec. arXiv. http://arxiv.org/abs/2107.03312v1
Chicago
Zeghidour, N., A. Luebs, A. Omran, J. Skoglund, and M. Tagliasacchi. 2021. “SoundStream: An End-to-End Neural Audio Codec”. arXiv. http://arxiv.org/abs/2107.03312v1.
Harvard
Zeghidour, N. et al. (2021) “SoundStream: An End-to-End Neural Audio Codec”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2107.03312v1.
Vancouver
1. Zeghidour N, Luebs A, Omran A, Skoglund J, Tagliasacchi M (2021) SoundStream: An End-to-End Neural Audio Codec. arXiv

BibTeX

@article{zeghidour2021soundstream,
  title = {SoundStream: An End-to-End Neural Audio Codec},
  author = {Zeghidour, Neil and Luebs, Alejandro and Omran, Ahmed and Skoglund, Jan and Tagliasacchi, Marco},
  year = {2021},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2107.03312v1},
  eprint = {2107.03312}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF