High Fidelity Neural Audio Compression

Alexandre D'efossezJade CopetGabriel SynnaeveYossi Adi

article2022Trans. Mach. Learn. Res.1,327 citations

Introduces a real-time, streaming neural audio codec that achieves state-of-the-art fidelity on monophonic and stereophonic audio by combining multiscale adversarial training, an innovative loss-balancing mechanism, and Transformer-based entropy coding.

Listen

Streaming media accounts for the vast majority of global internet traffic, driving a critical need for efficient audio compression that minimizes network bandwidth without degrading sound quality. Traditional codecs deliver acceptable fidelity at moderate bitrates but degrade significantly at very low bitrates, especially on complex audio such as music. The article evaluates EnCodec, a real-time, high-fidelity neural audio codec designed to compress speech and music across diverse bitrates and sample rates. It aims to demonstrate that deep-learning-based compression can achieve superior acoustic quality over existing industry standards while remaining computationally efficient.

The authors developed an end-to-end convolutional encoder-decoder network combined with residual vector quantization to compress audio into compact discrete representations. To refine training, they introduced a multi-scale spectrogram adversarial loss and a novel gradient balancer mechanism that stabilizes multi-objective optimization by decoupling hyperparameter tuning from loss scales. Optionally, a lightweight Transformer language model was applied for entropy coding to further compress the bitstream. The system was trained on thousands of hours of speech, music, and environmental sounds, and rigorously benchmarked against standard codecs like Opus and EVS as well as recent neural baselines via both objective metrics and subjective human listening evaluations.

The findings establish that EnCodec consistently outperforms standard and neural baselines across all evaluated settings. At 3 kilobits per second, EnCodec achieved higher perceptual quality scores than Opus at 12 kilobits per second and Lyra-v2 at 6 kilobits per second. In stereo music compression at 48 kilohertz, EnCodec at 6 kilobits per second matched the quality of MP3 compression at 64 kilobits per second, achieving a 256-to-1 compression ratio. Additionally, integrating the Transformer language model reduced the required bandwidth by 25% to 40% without perceptual degradation. Computationally, the streaming model operated roughly ten times faster than real time on a single standard computer processing core with an initial algorithmic latency of only 13.3 milliseconds.

These results demonstrate that high-fidelity audio transmission is achievable at unprecedentedly low bitrates, which can substantially reduce infrastructure and distribution costs for music streaming, teleconferencing, and telephony. Operating effectively at very low bitrates makes real-time communication feasible in poor connectivity environments without compromising user experience. For deployment, decision-makers should consider the streamable model for low-latency interactive applications, whereas the non-streamable setup paired with entropy coding provides maximum data savings for offline archiving and on-demand streaming.

Organizations should evaluate pilot implementations of the open-source code for low-bandwidth communication channels and streaming services. Before deploying entropy coding in strict real-time systems, teams must account for modest latency increases and evaluate floating-point precision across varied hardware to avoid decoding discrepancies. Furthermore, while the 24-kilohertz model runs easily in real time on a single CPU, high-resolution 48-kilohertz stereo processing with entropy coding is currently slower than real time on a single CPU core, indicating that hardware acceleration or code optimization is required before production rollout in live high-resolution streaming scenarios.

  • Paper: SoundStream: An End-to-End Neural Audio Codec, Neil Zeghidour et al. (2021). SoundStream establishes the core neural audio codec paradigm—combining a fully convolutional encoder-decoder, residual vector quantization, and adversarial/spectral losses—upon which EnCodec directly builds and refines.
  • Paper: Neural Discrete Representation Learning, Aäron van den Oord et al. (2017). This foundational paper introduces vector-quantized variational autoencoders (VQ-VAE) and straight-through gradient estimation, providing the underlying discrete representation technique used in neural audio quantization.
  • Paper: HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis, Jungil Kong et al. (2020). HiFi-GAN introduces multi-scale and multi-period discriminator formulations along with feature matching losses, serving as a primary foundation for high-fidelity neural audio synthesis and adversarial loss design.
  • Paper: Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers, Chengyi Wang et al. (2023). VALL-E directly leverages the discrete acoustic tokens from neural audio codecs like EnCodec to reformulate zero-shot text-to-speech synthesis as a conditional language modeling problem.
  • Paper: Qwen3-TTS Technical Report, Hangrui Hu et al. (2026). Qwen3-TTS extends the neural audio tokenization and multi-stage codec generation framework to multilingual, real-time zero-shot voice cloning and dual-track speech synthesis.
Cover for High Fidelity Neural Audio Compression

Abstract

We introduce a state-of-the-art real-time, high-fidelity, audio codec leveraging neural networks. It consists in a streaming encoder-decoder architecture with quantized latent space trained in an end-to-end fashion. We simplify and speed-up the training by using a single multiscale spectrogram adversary that efficiently reduces artifacts and produce high-quality samples. We introduce a novel loss balancer mechanism to stabilize training: the weight of a loss now defines the fraction of the overall gradient it should represent, thus decoupling the choice of this hyper-parameter from the typical scale of the loss. Finally, we study how lightweight Transformer models can be used to further compress the obtained representation by up to 40%, while staying faster than real time. We provide a detailed description of the key design choices of the proposed model including: training objective, architectural changes and a study of various perceptual loss functions. We present an extensive subjective evaluation (MUSHRA tests) together with an ablation study for a range of bandwidths and audio domains, including speech, noisy-reverberant speech, and music. Our approach is superior to the baselines methods across all evaluated settings, considering both 24 kHz monophonic and 48 kHz stereophonic audio. Code and models are available at this http URL.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Model
  • 3.1 Encoder & Decoder Architecture
  • 3.2 Residual Vector Quantization
  • 3.3 Language Modeling and Entropy Coding
  • 3.4 Training objective
  • 4 Experiments and Results
  • 4.1 Dataset
  • 4.2 Baselines
  • 4.3 Evaluation Methods
  • 4.4 Training
  • 4.5 Results
  • 4.5.1 Ablation study
  • 4.5.2 Stereo Evaluation
  • 4.6 Latency and computation time
  • 5 Conclusion
  • References
  • A Appendix
  • A.1 Experimental details
  • A.2 Alternative quantizers
  • A.2.1 DiffQ Quantizer
  • A.2.2 Gumbel softmax quantizer
  • A.3 Additional Results
  • A.4 Societal impact

Knowls

  1. Knowl 1 — EnCodec Streaming Neural Audio Codec Architecture

    model/method

    EnCodec is an end-to-end neural audio compression architecture comprising an encoder EE, a Residual Vector Quantizer (RVQ) QQ, and a decoder GG.

    Given an input audio signal x∈[−1,1]Ca×Tx \in [-1, 1]^{C_a \times T} with CaC_a audio channels and TT samples:

    1. Encoder (EE): Consists of a 1D convolution with C=32C=32 initial channels and kernel size 7, followed by B=4B=4 convolutional blocks. Each block contains a single residual unit (two convolutions of kernel size 3 with a skip connection) followed by a strided downsampling convolution with stride S∈{2,4,5,8}S \in \{2, 4, 5, 8\} (total downsampling factor 320) and kernel size K=2SK = 2S. Channels double upon downsampling. The blocks are followed by a 2-layer LSTM for sequence modeling and a final 1D convolution with kernel size 7 and D=128D=128 output channels. At 24 kHz audio sampling rate, the encoder produces 75 latent steps per second; at 48 kHz, it produces 150 latent steps per second.
    2. Residual Vector Quantization (QQ): Quantizes the continuous latent representation z∈RB×D×T′z \in \mathbb{R}^{B \times D \times T'} into NqN_q discrete codebook indices using NqN_q cascaded vector quantizers (up to 32 codebooks for 24 kHz and 16 codebooks for 48 kHz, each containing 1024 entries of 10 bits). Each stage quantizes the residual error of the preceding stages.
    3. Decoder (GG): Symmetrically mirrors the encoder using transposed convolutions in reverse stride order to reconstruct the waveform x^\hat{x}.

    The architecture operates in two modes:

    • Streamable Mode: Applies causal padding (all K−SK-S padding placed before the first time step), replaces layer normalization with weight normalization, and processes transposed convolutions such that the first ss steps are output immediately while the remaining ss steps are buffered for the next frame. This achieves an initial algorithmic latency of 320 samples (13.3 ms at 24 kHz).
    • Non-streamable Mode: Applies centered padding ((K−S)/2(K-S)/2 before and after), splits audio into 1-second chunks with 10 ms overlap, and normalizes each chunk using layer normalization across both channel and time dimensions.
  2. Knowl 2 — Gradient Loss Balancer for Multi-Objective Training

    model/method

    The loss balancer is a training stabilization mechanism designed to decouple hyperparameter weights from the intrinsic gradient scales of multiple competing loss functions.

    Let (li)i=1M(l_i)_{i=1}^M be a set of loss terms depending on the model output x^\hat{x}, and let gi=∂li∂x^g_i = \frac{\partial l_i}{\partial \hat{x}} denote the gradient of loss lil_i with respect to x^\hat{x}. Let ⟨∥gi∥2⟩β\langle \|g_i\|_2 \rangle_\beta be the exponential moving average of the Euclidean norm of gig_i computed across training batches with momentum parameter β=0.999\beta = 0.999. Given target loss weights (λi)i=1M(\lambda_i)_{i=1}^M and a reference gradient norm R=1R = 1, the rescaled gradient g~i\tilde{g}_i backpropagated for loss lil_i is:

    g~i=R⋅λi∑jλj⋅gi⟨∥gi∥2⟩β\tilde{g}_i = R \cdot \frac{\lambda_i}{\sum_j \lambda_j} \cdot \frac{g_i}{\langle \|g_i\|_2 \rangle_\beta}

    The total gradient injected into the network during backpropagation is ∑ig~i\sum_i \tilde{g}_i. When normalized such that ∑iλi=1\sum_i \lambda_i = 1, each coefficient λi\lambda_i directly defines the exact fraction of the total model gradient contributed by loss lil_i, preventing discriminators or individual reconstruction losses from dominating optimization regardless of the raw loss magnitudes.

  3. Knowl 3 — Multi-Scale STFT Discriminator (MS-STFTD)

    model/method

    The Multi-Scale Short-Time Fourier Transform Discriminator (MS-STFTD) is a non-waveform, spectrogram-based adversarial architecture that evaluates audio fidelity across multiple time-frequency resolutions.

    MS-STFTD consists of 5 sub-networks operating on complex-valued STFT representations with window sizes w∈{2048,1024,512,256,128}w \in \{2048, 1024, 512, 256, 128\} and hop lengths h=w/4h = w/4 for 24 kHz audio (window lengths are doubled for 48 kHz audio). The real and imaginary STFT components are concatenated along the channel dimension.

    Each sub-network architecture comprises:

    1. An initial 2D convolutional layer with 32 channels, kernel size 3×83 \times 8, stride (1,2)(1, 2), and dilation (1,1)(1, 1).
    2. Successive 2D convolutional blocks with time dilation rates of 1, 2, and 4, and a frequency stride of 2.
    3. A final 2D convolutional layer with kernel size 3×33 \times 3 and stride (1,1)(1, 1) outputting scalar classification logits.

    All layers utilize LeakyReLU activations and weight normalization. For multi-channel (stereo) audio, left and right channels are evaluated independently by the discriminators.

  4. Knowl 4 — EnCodec Training Objectives and Relative Feature Matching Loss

    equation

    The EnCodec generator GG is trained to minimize a composite loss function comprising time-domain reconstruction, multi-scale spectral reconstruction, adversarial hinge loss, relative feature matching, and quantizer commitment:

    LG=λtlt(x,x^)+λflf(x,x^)+λglg(x^)+λfeatlfeat(x,x^)+λwlwL_G = \lambda_t l_t(x, \hat{x}) + \lambda_f l_f(x, \hat{x}) + \lambda_g l_g(\hat{x}) + \lambda_{feat} l_{feat}(x, \hat{x}) + \lambda_w l_w

    where xx is the ground-truth audio and x^\hat{x} is the reconstructed audio.

    1. Time-Domain Loss: lt(x,x^)=∥x−x^∥1l_t(x, \hat{x}) = \|x - \hat{x}\|_1

    2. Frequency-Domain Loss: lf(x,x^)=1∣e∣∑i∈e(∥Si(x)−Si(x^)∥1+∥Si(x)−Si(x^)∥2)l_f(x, \hat{x}) = \frac{1}{|e|} \sum_{i \in e} \left( \|S_i(x) - S_i(\hat{x})\|_1 + \|S_i(x) - S_i(\hat{x})\|_2 \right) where SiS_i is a 64-bin normalized mel-spectrogram with STFT window size 2i2^i and hop length 2i/42^i / 4 over scales e={5,6,7,8,9,10,11}e = \{5, 6, 7, 8, 9, 10, 11\}.

    3. Adversarial Generator Loss: lg(x^)=1K∑k=1Kmax⁡(0,1−Dk(x^))l_g(\hat{x}) = \frac{1}{K} \sum_{k=1}^K \max(0, 1 - D_k(\hat{x})) where KK is the number of discriminators and DkD_k is the kk-th discriminator.

    4. Relative Feature Matching Loss: lfeat(x,x^)=1KL∑k=1K∑l=1L∥Dkl(x)−Dkl(x^)∥1mean(∥Dkl(x)∥1)l_{feat}(x, \hat{x}) = \frac{1}{K L} \sum_{k=1}^K \sum_{l=1}^L \frac{\|D_k^l(x) - D_k^l(\hat{x})\|_1}{\mathrm{mean}(\|D_k^l(x)\|_1)} where DklD_k^l denotes the activation output of the ll-th layer of discriminator DkD_k, and the mean in the denominator is computed across all feature dimensions.

    5. RVQ Commitment Loss: lw=∑c=1C∥zc−qc(zc)∥22l_w = \sum_{c=1}^C \|z_c - q_c(z_c)\|_2^2 where zcz_c is the residual input at RVQ stage cc, qc(zc)q_c(z_c) is the nearest codebook centroid, and CC is the active number of codebooks.

  5. Knowl 5 — Language Model Entropy Coding for RVQ Latents

    model/method

    To compress the discrete indices produced by Residual Vector Quantization (RVQ) beyond their fixed-rate representation, an autoregressive Transformer language model is applied prior to arithmetic encoding.

    The Transformer architecture consists of 5 layers, 8 attention heads, a latent dimension of 200, a feed-forward inner dimension of 800, no dropout, and a causal receptive field of 3.5 seconds. For a sequence of discrete codebook indices with shape [Nq,T][N_q, T], where NqN_q is the number of codebooks and TT is the number of time frames:

    1. At time step tt, the discrete codes from the preceding frame t−1t-1 across all NqN_q codebooks are converted into vectors via learned embedding tables (one table per codebook) and summed into a single continuous representation.
    2. The Transformer processes the continuous sequence and passes its output through NqN_q independent linear heads of dimension 1024 (matching codebook cardinality) to predict the unnormalized log-probabilities for all NqN_q codebooks at time step tt simultaneously, deliberately neglecting intra-frame cross-codebook mutual information to enable faster-than-real-time inference.
    3. A range-based arithmetic coder encodes the tokens using predicted probabilities rounded to a fixed precision of 10−610^{-6} with a total range width of 2242^{24} and a minimum range width of 2.

    This entropy coding stage reduces overall transmission bitrates by 25% to 40% without introducing any waveform distortion.

  6. Knowl 6 — MUSHRA Subjective Evaluation on 24 kHz Monophonic Audio

    data/table

    Subjective listening tests following the MUSHRA protocol (scale 0–100) were conducted on 24 kHz monophonic audio across clean speech (DNS Challenge 4), noisy speech (DNS + FSD50K), and music (Jamendo Set-1 and proprietary Set-2), comparing streamable EnCodec against traditional codecs (Opus, EVS) and neural codecs (Lyra-v2).

    Model Bandwidth Entropy Coded Clean Speech Noisy Speech Music Set-1 Music Set-2
    Reference - - 95.5 1.6 93.9 1.8 93.2 2.5 97.1 1.3
    Opus 6.0 kbps - 30.1 2.8 19.1 5.9 20.6 5.8 17.9 5.3
    Opus 12.0 kbps - 76.5 2.3 61.9 2.1 77.8 3.2 65.4 2.7
    EVS 9.6 kbps - 84.4 2.5 80.0 2.4 89.9 2.3 87.7 2.3
    Lyra-v2 3.0 kbps - 53.1 1.9 52.0 4.7 69.3 3.3 42.3 3.5
    Lyra-v2 6.0 kbps - 66.2 2.9 59.9 3.3 75.7 2.6 48.6 2.1
    EnCodec 1.5 kbps 0.9 kbps 49.2 2.4 41.3 3.6 68.2 2.2 66.5 2.3
    EnCodec 3.0 kbps 1.9 kbps 67.0 1.5 62.5 2.3 89.6 3.1 87.8 2.9
    EnCodec 6.0 kbps 4.1 kbps 83.1 2.7 69.4 2.3 92.9 1.8 91.3 2.1
    EnCodec 12.0 kbps 8.9 kbps 90.6 2.6 80.1 2.5 91.8 2.5 92.9 1.2

    EnCodec at 3.0 kbps (1.9 kbps with entropy coding) outperforms Lyra-v2 at 6.0 kbps across speech and music and matches or exceeds Opus at 12.0 kbps on music. At 12.0 kbps, EnCodec approaches the reference score across speech and music domains.

  7. Knowl 7 — MUSHRA Subjective Evaluation on 48 kHz Stereophonic Audio

    data/table

    Subjective MUSHRA evaluation on 48 kHz stereophonic music comparing EnCodec against MP3 and Opus across extreme compression ratios:

    Model Bandwidth Entropy Coded Compression MUSHRA
    Reference - - 1× 95.1 1.8
    MP3 64 kbps - 24× 82.7 3.2
    Opus 6 kbps - 256× 17.7 5.9
    Opus 24 kbps - 64× 82.9 3.7
    EnCodec 6 kbps 4.2 kbps 256× 82.9 2.4
    EnCodec 12 kbps 8.9 kbps 128× 88.0 2.7
    EnCodec 24 kbps 19.4 kbps 64× 87.5 2.6

    At 6 kbps (256×256\times compression, or 4.2 kbps with entropy coding), EnCodec achieves a MUSHRA score of 82.9±2.482.9 \pm 2.4, matching MP3 at 64 kbps (24×24\times compression, 82.7±3.282.7 \pm 3.2) and Opus at 24 kbps (64×64\times compression, 82.9±3.782.9 \pm 3.7), while Opus at 6 kbps collapses to 17.7±5.917.7 \pm 5.9.

  8. Knowl 8 — Ablation Study of Audio Discriminator Architectures

    data/table

    An evaluation of objective metrics (SI-SNR, ViSQOL) and subjective MUSHRA scores comparing different discriminator combinations during training of EnCodec:

    Discriminator Setup SI-SNR ViSQOL MUSHRA
    MSD + Mono-STFT 5.99 4.22 62.91 2.62
    MPD 7.35 4.24 60.7 2.8
    MS-STFT + MPD 6.55 4.34 79.0 1.9
    MS-STFT 6.67 4.35 77.5 1.8

    Using solely the Multi-Scale STFT Discriminator (MS-STFT) yields a MUSHRA score of 77.5±1.877.5 \pm 1.8 and ViSQOL of 4.35, substantially outperforming the baseline combination of Multi-Scale Discriminator and single STFT discriminator (MSD + Mono-STFT, 62.91±2.6262.91 \pm 2.62) and waveform Multi-Period Discriminator alone (MPD, 60.7±2.860.7 \pm 2.8). Adding MPD to MS-STFT provides a marginal gain (79.0±1.979.0 \pm 1.9), indicating that multi-scale spectral discrimination alone is sufficient for high perceptual fidelity while simplifying the training pipeline.

  9. Knowl 9 — Initial Latency and Real-Time Factor (RTF)

    data/table

    Computational profiling for EnCodec evaluated at 6 kbps on a single thread of an Intel Core CPU (MacBook Pro 2019). The Real-Time Factor (RTF) is defined as the audio duration divided by the processing time, where RTF>1\text{RTF} > 1 represents faster-than-real-time operation.

    Model Latency Enc. RTF Dec. RTF Enc. + EC RTF Dec. + EC RTF
    Lyra v2 (32 kHz) - 27.4 67.2 - -
    EnCodec 24 kHz 13 ms 9.8 10.4 1.6 1.6
    EnCodec 48 kHz 1 s 6.8 5.1 0.68 0.66

    Without entropy coding (EC), the 24 kHz streaming EnCodec model operates approximately 10×10\times faster than real time (extRTF=9.8 ext{RTF} = 9.8 for encoder, 10.410.4 for decoder) with an algorithmic latency of 13.3 ms. With the Transformer language model and arithmetic entropy coder enabled, 24 kHz EnCodec maintains real-time operation (RTF=1.6\text{RTF} = 1.6) with an additional 13 ms buffering latency. The 48 kHz non-streaming model exhibits an initial latency of 1.0 s due to chunk normalization.

  10. Knowl 10 — Differentiable Latent Quantization via Pseudo Quantization Noise (DiffQ)

    model/method

    DiffQ is an alternative differentiable latent quantization mechanism using additive pseudo-quantization noise during training to optimize per-channel bit allocations.

    Let z∈RDz \in \mathbb{R}^D be the latent vector produced by the encoder, with mean mm and standard deviation σ\sigma computed over the batch and time dimensions. A learnable parameter B∈RDB \in \mathbb{R}^D represents the bit allocation for each dimension, parameterized as B=Bmax⋅sigmoid(αv)B = B_{max} \cdot \mathrm{sigmoid}(\alpha v) with Bmax=15B_{max} = 15 and α=5\alpha = 5.

    1. Training Phase: Quantization is simulated by injecting uniform noise U[−1,1]U[-1, 1] scaled by the bit depth: zq,train=clamp(z,m−Lσ,m+Lσ)+L⋅σ⋅U[−1,1]2Bz_{q,train} = \mathrm{clamp}(z, m - L\sigma, m + L\sigma) + L \cdot \sigma \cdot \frac{U[-1, 1]}{2^B} where L=3L = 3 sets the dynamic range clamping threshold (preserving 99.5%99.5\% of Gaussian distributed latents). To enforce true sparsity as B→0B \to 0, zz is scaled by min⁡(B,1)\min(B, 1) while noise is scaled by min⁡(B,1)\sqrt{\min(B, 1)}. Bandwidth is regularized via penalty λwdiffq⋅max⁡(0,wdiffq−wtarget)\lambda_{w_{diffq}} \cdot \max(0, w_{diffq} - w_{target}) with estimated bitrate wdiffq=T′d∑i=1DB(i)w_{diffq} = \frac{T'}{d} \sum_{i=1}^D B^{(i)}.

    2. Inference Phase: Latents are quantized deterministically to NB=round(2B)N_B = \mathrm{round}(2^B) levels via normalized coordinates u=clamp(L+σ−1(z−m)2L,0,1)u = \mathrm{clamp}\left(\frac{L + \sigma^{-1}(z-m)}{2L}, 0, 1\right): i=min⁡(⌊NB⋅u⌋,NB−1)∈{0,…,NB−1}i = \min(\lfloor N_B \cdot u \rfloor, N_B - 1) \in \{0, \dots, N_B - 1\} zq=m+Lσ(2i+0.5NB−1)z_q = m + L\sigma \left( 2\frac{i + 0.5}{N_B} - 1 \right)

  11. Knowl 11 — Training Stability Impact of Gradient Balancer under Adversarial Loss Weight Scaling

    empirical result

    Empirical evaluation of the gradient loss balancer demonstrates that it prevents training divergence across wide variations in adversarial and reconstruction loss hyperparameters (λt,λf,λg,λfeat)(\lambda_t, \lambda_f, \lambda_g, \lambda_{feat}).

    When training EnCodec without the loss balancer on music data:

    • Increasing the adversarial generator loss weight λg\lambda_g from 1 to 100 causes severe generator collapse, resulting in SI-SNR dropping from 6.16 dB6.16\text{ dB} to −35.83 dB-35.83\text{ dB} and ViSQOL dropping from 3.893.89 to 2.822.82.
    • Conversely, when the gradient balancer is active with the same parameters (λt=1,λf=2,λg=100,λfeat=1)(\lambda_t=1, \lambda_f=2, \lambda_g=100, \lambda_{feat}=1), the model maintains stable training with SI-SNR of 8.41 dB8.41\text{ dB} and ViSQOL of 4.054.05.
    • Across all tested loss weight permutations (with λg\lambda_g varying from 1 to 100 and λt\lambda_t from 1 to 10), the balancer maintains SI-SNR in the range [8.41,10.72] dB[8.41, 10.72]\text{ dB} and ViSQOL in the range [3.62,4.17][3.62, 4.17], demonstrating robustness against hyperparameter miscalibration.

Coverage note — None was omitted; all primary architectural components, losses, entropy coding models, alternative quantizers (DiffQ, Gumbel-Softmax), and core empirical results across mono, stereo, ablations, and latency benchmarks are represented.

References

  1. 1.Pavel Andreev, Aibek Alanov, Oleg Ivanov, and Dmitry Vetrov. Hifi++: a unified framework for neural vocoding, bandwidth extension and speech enhancement. arXiv preprint arXiv:2203.13086, 2022.
  2. 2.Rosana Ardila, Megan Branson, Kelly Davis, Michael Henretty, Michael Kohler, Josh Meyer, Reuben Morais, Lindsay Saunders, Francis M Tyers, and Gregor Weber. Common voice: A massively-multilingual speech corpus. arXiv preprint arXiv:1912.06670, 2019.
  3. 3.Bishnu S Atal and Suzanne L Hanauer. Speech analysis and synthesis by linear prediction of the speech wave. The journal of the acoustical society of America, 50(2B):637–655, 1971.
  4. 4.Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton. Layer normalization. arXiv preprint arXiv:1607.06450, 2016.
  5. 5.Johannes Ballé, Valero Laparra, and Eero P Simoncelli. End-to-end optimized image compression. In ICLR, 2017.
  6. 6.Johannes Ballé, Nick Johnston, and David Minnen. Integer networks for data compression with latent-variable models. In International Conference on Learning Representations, 2018.
  7. 7.Yoshua Bengio, Nicholas Léonard, and Aaron Courville. Estimating or propagating gradients through stochastic neurons for conditional computation. arXiv preprint arXiv:1308.3432, 2013.
  8. 8.Dipesh Bhagat, Ninad Bhatt, and Yogeshwar Kosta. Adaptive multi-rate wideband speech codec based on celp algorithm: architectural study, implementation & performance analysis. In 2012 International Conference on Communication Systems and Network Technologies, pp. 547–551. IEEE, 2012.
  9. 9.Dmitry Bogdanov, Minz Won, Philip Tovstogan, Alastair Porter, and Xavier Serra. The mtg-jamendo dataset for automatic music tagging. In Machine Learning for Music Discovery Workshop, International Conference on Machine Learning (ICML 2019), Long Beach, CA, United States, 2019. URL http://hdl.handle.net/10230/42015.
  10. 10.Shlomo E Chazan, Lior Wolf, Eliya Nachmani, and Yossi Adi. Single channel voice separation for unknown number of speakers under reverberant and noisy settings. In ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 3730–3734. IEEE, 2021.
  11. 11.Michael Chinen, Felicia SC Lim, Jan Skoglund, Nikita Gureev, Feargus O’Gorman, and Andrew Hines. Visqol v3: An open source production ready objective speech and audio metric. In 2020 twelfth international conference on quality of multimedia experience (QoMEX), pp. 1–6. IEEE, 2020.
  12. 12.Cisco. Global - 2021 forecast highlights - cisco. https://www.cisco.com/c/dam/m/en_us/solutions/service-provider/vni-forecast-highlights/pdf/Global_2021_Forecast_Highlights.pdf, 2021.
  13. 13.Djork-Arné Clevert, Thomas Unterthiner, and Sepp Hochreiter. Fast and accurate deep network learning by exponential linear units (elus). arXiv preprint arXiv:1511.07289, 2015.
  14. 14.Alexandre Défossez, Nicolas Usunier, Léon Bottou, and Francis Bach. Music source separation in the waveform domain. arXiv preprint arXiv:1911.13254, 2019.
  15. 15.Alexandre Defossez, Gabriel Synnaeve, and Yossi Adi. Real time speech enhancement in the waveform domain. arXiv preprint arXiv:2006.12847, 2020.
  16. 16.Alexandre Défossez, Yossi Adi, and Gabriel Synnaeve. Differentiable model compression via pseudo quantization noise. arXiv preprint arXiv:2104.09987, 2021.
  17. 17.Prafulla Dhariwal, Heewoo Jun, Christine Payne, Jong Wook Kim, Alec Radford, and Ilya Sutskever. Jukebox: A generative model for music. arXiv preprint arXiv:2005.00341, 2020.
  18. 18.Sander Dieleman, Aaron van den Oord, and Karen Simonyan. The challenge of realistic music generation: modelling raw audio at scale. Advances in Neural Information Processing Systems, 31, 2018.
  19. 19.Martin Dietz, Markus Multrus, Vaclav Eksler, Vladimir Malenovsky, Erik Norvell, Harald Pobloth, Lei Miao, Zhe Wang, Lasse Laaksonen, Adriana Vasilache, et al. Overview of the evs codec architecture. In 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 5698–5702. IEEE, 2015.
  20. 20.Harishchandra Dubey, Vishak Gopal, Ross Cutler, Sergiy Matusevych, Sebastian Braun, Emre Sefik Eskimez, Manthan Thakker, Takuya Yoshioka, Hannes Gamper, and Robert Aichner. Icassp 2022 deep noise suppression challenge. In ICASSP, 2022.
  21. 21.Eduardo Fonseca, Xavier Favory, Jordi Pons, Frederic Font, and Xavier Serra. Fsd50k: an open dataset of human-labeled sound events. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 30: 829–852, 2021.
  22. 22.Cristina Gârbacea, Aäron van den Oord, Yazhe Li, Felicia SC Lim, Alejandro Luebs, Oriol Vinyals, and Thomas C Walters. Low bit-rate speech coding with vq-vae and a wavenet decoder. In ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 735–739. IEEE, 2019.
  23. 23.Jort F Gemmeke, Daniel PW Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R Channing Moore, Manoj Plakal, and Marvin Ritter. Audio set: An ontology and human-labeled dataset for audio events. In 2017 IEEE international conference on acoustics, speech and signal processing (ICASSP), pp. 776–780. IEEE, 2017.
  24. 24.Karan Goel, Albert Gu, Chris Donahue, and Christopher Ré. It’s raw! audio generation with state-space models. arXiv preprint arXiv:2202.09729, 2022.
  25. 25.Robert Gray. Vector quantization. IEEE Assp Magazine, 1(2):4–29, 1984.
  26. 26.D Griffin and Jae Lim. A new model-based speech analysis/synthesis system. In ICASSP, 1985.
  27. 27.Alexey Gritsenko, Tim Salimans, Rianne van den Berg, Jasper Snoek, and Nal Kalchbrenner. A spectral energy distance for parallel speech synthesis. Advances in Neural Information Processing Systems, 33: 13062–13072, 2020.
  28. 28.Andrew Hines, Jan Skoglund, Anil Kokaram, and Naomi Harte. Visqol: The virtual speech quality objective listener. In IWAENC 2012; International Workshop on Acoustic Signal Enhancement, pp. 1–4. VDE, 2012.
  29. 29.Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed. Hubert: Self-supervised speech representation learning by masked prediction of hidden units. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 29:3451–3460, 2021.
  30. 30.Sergey Ioffe and Christian Szegedy. Batch normalization: Accelerating deep network training by reducing internal covariate shift. Technical Report 1502.03167, arXiv, 2015.
  31. 31.Eric Jang, Shixiang Gu, and Ben Poole. Categorical reparameterization with gumbel-softmax. In ICLR, 2017.
  32. 32.Tejas Jayashankar, Thilo Koehler, Kaustubh Kalgaonkar, Zhiping Xiu, Jilong Wu, Ju Lin, Prabhav Agrawal, and Qing He. Architecture for variable bitrate neural speech codec with configurable computation complexity. In ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 861–865. IEEE, 2022.
  33. 33.Xue Jiang, Xiulian Peng, Chengyu Zheng, Huaying Xue, Yuan Zhang, and Yan Lu. End-to-end neural speech coding for real-time communications. In ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 866–870. IEEE, 2022.
  34. 34.Biing-Hwang Juang and A Gray. Multiple stage vector quantization for speech coding. In ICASSP’82. IEEE International Conference on Acoustics, Speech, and Signal Processing, volume 7, pp. 597–600. IEEE, 1982.
  35. 35.Nal Kalchbrenner et al. Efficient Neural Audio Synthesis. In ICML, 2018.
  36. 36.Eugene Kharitonov, Ann Lee, Adam Polyak, Yossi Adi, Jade Copet, Kushal Lakhotia, Tu-Anh Nguyen, Morgane Rivière, Abdelrahman Mohamed, Emmanuel Dupoux, et al. Text-free prosody-aware generative spoken language modeling. arXiv preprint arXiv:2109.03264, 2021.
  37. 37.W Bastiaan Kleijn, Felicia SC Lim, Alejandro Luebs, Jan Skoglund, Florian Stimberg, Quan Wang, and Thomas C Walters. Wavenet based low rate speech coding. In 2018 IEEE international conference on acoustics, speech and signal processing (ICASSP), pp. 676–680. IEEE, 2018.
  38. 38.W Bastiaan Kleijn, Andrew Storus, Michael Chinen, Tom Denton, Felicia SC Lim, Alejandro Luebs, Jan Skoglund, and Hengchin Yeh. Generative speech coding with predictive variance regularization. In ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 6478–6482. IEEE, 2021.
  39. 39.Jungil Kong, Jaehyeon Kim, and Jaekyoung Bae. Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis. Advances in Neural Information Processing Systems, 33:17022–17033, 2020.
  40. 40.Felix Kreuk, Adam Polyak, Jade Copet, Eugene Kharitonov, Tu-Anh Nguyen, Morgane Rivière, Wei-Ning Hsu, Abdelrahman Mohamed, Emmanuel Dupoux, and Yossi Adi. Textless speech emotion conversion using decomposed and discrete representations. arXiv preprint arXiv:2111.07402, 2021.
  41. 41.Kundan Kumar, Rithesh Kumar, Thibault de Boissiere, Lucas Gestin, Wei Zhen Teoh, Jose Sotelo, Alexandre de Brébisson, Yoshua Bengio, and Aaron C Courville. Melgan: Generative adversarial networks for conditional waveform synthesis. Advances in neural information processing systems, 32, 2019.
  42. 42.Kushal Lakhotia, Eugene Kharitonov, Wei-Ning Hsu, Yossi Adi, Adam Polyak, Benjamin Bolte, Tu-Anh Nguyen, Jade Copet, Alexei Baevski, Abdelrahman Mohamed, et al. On generative spoken language modeling from raw audio. Transactions of the Association for Computational Linguistics, 9:1336–1354, 2021.
  43. 43.Ann Lee, Peng-Jen Chen, Changhan Wang, Jiatao Gu, Xutai Ma, Adam Polyak, Yossi Adi, Qing He, Yun Tang, Juan Pino, et al. Direct speech-to-speech translation with discrete units. arXiv preprint arXiv:2107.05604, 2021a.
  44. 44.Ann Lee, Hongyu Gong, Paul-Ambroise Duquenne, Holger Schwenk, Peng-Jen Chen, Changhan Wang, Sravya Popuri, Juan Pino, Jiatao Gu, and Wei-Ning Hsu. Textless speech-to-speech translation on real data. arXiv preprint arXiv:2112.08352, 2021b.
  45. 45.Shuyang Li, Huanru Henry Mao, and Julian McAuley. Variable bitrate discrete neural representations via causal self-attention. In 2nd Pre-registration workshop (NeurIPS 2021), Remote.
  46. 46.Yunpeng Li, Marco Tagliasacchi, Oleg Rybakov, Victor Ungureanu, and Dominik Roblek. Real-time speech frequency bandwidth extension. In ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 691–695. IEEE, 2021.
  47. 47.Felicia SC Lim, W Bastiaan Kleijn, Michael Chinen, and Jan Skoglund. Robust low rate speech coding based on cloned networks and wavenet. In ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 6769–6773. IEEE, 2020.
  48. 48.Ju Lin, Kaustubh Kalgaonkar, Qing He, and Xin Lei. Speech enhancement for low bit rate speech codec. In ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 7777–7781. IEEE, 2022.
  49. 49.Yi Luo and Nima Mesgarani. Conv-tasnet: Surpassing ideal time–frequency magnitude masking for speech separation. IEEE/ACM transactions on audio, speech, and language processing, 27(8):1256–1266, 2019.
  50. 50.Alan McCree, Kwan Truong, E Bryan George, Thomas P Barnwell, and Vishu Viswanathan. A 2.4 kbit/s melp coder candidate for the new us federal standard. In ICASSP, 1996.
  51. 51.Shigeo Morishima, H Harashima, and Y Katayama. Speech coding based on a multi-layer neural network. In IEEE International Conference on Communications, Including Supercomm Technical Sessions, pp. 429–433. IEEE, 1990.
  52. 52.Eliya Nachmani, Yossi Adi, and Lior Wolf. Voice separation with an unknown number of multiple speakers. In International Conference on Machine Learning, pp. 7164–7175. PMLR, 2020.
  53. 53.Tu Anh Nguyen, Eugene Kharitonov, Jade Copet, Yossi Adi, Wei-Ning Hsu, Ali Elkahky, Paden Tomasello, Robin Algayres, Benoit Sagot, Abdelrahman Mohamed, et al. Generative spoken dialogue language modeling. arXiv preprint arXiv:2203.16502, 2022.
  54. 54.Ahmed Omran, Neil Zeghidour, Zalán Borsos, Félix de Chaumont Quitry, Malcolm Slaney, and Marco Tagliasacchi. Disentangling speech from surroundings in a neural audio codec. arXiv preprint arXiv:2203.15578, 2022.
  55. 55.Aaron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu. Wavenet: A generative model for raw audio. arXiv preprint arXiv:1609.03499, 2016.
  56. 56.Richard Clark Pasco. Source coding algorithms for fast data compression. PhD thesis, Stanford University CA, 1976.
  57. 57.Adam Polyak, Yossi Adi, Jade Copet, Eugene Kharitonov, Kushal Lakhotia, Wei-Ning Hsu, Abdelrahman Mohamed, and Emmanuel Dupoux. Speech resynthesis from discrete disentangled self-supervised representations. arXiv preprint arXiv:2104.00355, 2021.
  58. 58.Sravya Popuri, Peng-Jen Chen, Changhan Wang, Juan Pino, Yossi Adi, Jiatao Gu, Wei-Ning Hsu, and Ann Lee. Enhanced direct speech-to-speech translation using self-supervised pre-training and data augmentation. arXiv preprint arXiv:2204.02967, 2022.
  59. 59.Ali Razavi, Aaron Van den Oord, and Oriol Vinyals. Generating diverse high-fidelity images with vq-vae-2. Advances in neural information processing systems, 32, 2019.
  60. 60.Oren Rippel, Sanjay Nair, Carissa Lew, Steve Branson, Alexander G Anderson, and Lubomir Bourdev. Learned video compression. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 3454–3463, 2019.
  61. 61.Jorma Rissanen and Glen Langdon. Universal modeling and coding. IEEE Transactions on Information Theory, 27(1):12–23, 1981.
  62. 62.Tim Salimans and Durk P Kingma. Weight normalization: A simple reparameterization to accelerate training of deep neural networks. Advances in neural information processing systems, 29, 2016.
  63. 63.B Series. Method for the subjective assessment of intermediate quality level of audio systems. International Telecommunication Union Radiocommunication Assembly, 2014.
  64. 64.Jan Skoglund and Jean-Marc Valin. Improving opus low bit rate quality with neural speech synthesis. arXiv preprint arXiv:1905.04628, 2019.
  65. 65.Marco Tagliasacchi, Yunpeng Li, Karolis Misiunas, and Dominik Roblek. Seanet: A multi-modal speech enhancement network. arXiv preprint arXiv:2009.02095, 2020.
  66. 66.Jean-Marc Valin and Jan Skoglund. Lpcnet: Improving neural speech synthesis through linear prediction. In ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 5891–5895. IEEE, 2019a.
  67. 67.Jean-Marc Valin and Jan Skoglund. A real-time wideband neural vocoder at 1.6 kb/s using lpcnet. arXiv preprint arXiv:1903.12087, 2019b.
  68. 68.Jean-Marc Valin, Koen Vos, and Timothy Terriberry. Definition of the opus audio codec. IETF, September, 2, 2012.
  69. 69.Aaron Van Den Oord, Oriol Vinyals, et al. Neural discrete representation learning. Advances in neural information processing systems, 30, 2017.
  70. 70.A Vasuki and PT Vanathi. A review of vector quantization techniques. IEEE Potentials, 25(4):39–47, 2006.
  71. 71.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Proc. of Neural Information Processing Systems, 2017.
  72. 72.Bernard Widrow, Istvan Kollar, and Ming-Chang Liu. Statistical theory of quantization. IEEE Transactions on instrumentation and measurement, 45(2):353–361, 1996.
  73. 73.Ryuichi Yamamoto, Eunwoo Song, and Jae-Min Kim. Parallel wavegan: A fast waveform generation model based on generative adversarial networks with multi-resolution spectrogram. In ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 6199–6203. IEEE, 2020a.
  74. 74.Ryuichi Yamamoto, Eunwoo Song, and Jae-Min Kim. Parallel wavegan: A fast waveform generation model based on generative adversarial networks with multi-resolution spectrogram. In ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 6199–6203. IEEE, 2020b.
  75. 75.Jaeseong You, Dalhyun Kim, Gyuhyeon Nam, Geumbyeol Hwang, and Gyeongsu Chae. Gan vocoder: Multi-resolution discriminator is all you need. arXiv preprint arXiv:2103.05236, 2021.
  76. 76.Neil Zeghidour, Alejandro Luebs, Ahmed Omran, Jan Skoglund, and Marco Tagliasacchi. Soundstream: An end-to-end neural audio codec. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2021.

Citation

MLA
Défossez, A., et al. “High Fidelity Neural Audio Compression”. arXiv, 2022, http://arxiv.org/abs/2210.13438v1.
APA
Défossez, A., Copet, J., Synnaeve, G., & Adi, Y. (2022). High Fidelity Neural Audio Compression. arXiv. http://arxiv.org/abs/2210.13438v1
Chicago
Défossez, A., J. Copet, G. Synnaeve, and Y. Adi. 2022. “High Fidelity Neural Audio Compression”. arXiv. http://arxiv.org/abs/2210.13438v1.
Harvard
Défossez, A. et al. (2022) “High Fidelity Neural Audio Compression”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2210.13438v1.
Vancouver
1. Défossez A, Copet J, Synnaeve G, Adi Y (2022) High Fidelity Neural Audio Compression. arXiv

BibTeX

@article{defossez2022high,
  title = {High Fidelity Neural Audio Compression},
  author = {Défossez, Alexandre and Copet, Jade and Synnaeve, Gabriel and Adi, Yossi},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2210.13438v1},
  eprint = {2210.13438}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF