High Fidelity Neural Audio Compression
Alexandre D'efossezJade CopetGabriel SynnaeveYossi Adi
Introduces a real-time, streaming neural audio codec that achieves state-of-the-art fidelity on monophonic and stereophonic audio by combining multiscale adversarial training, an innovative loss-balancing mechanism, and Transformer-based entropy coding.
Streaming media accounts for the vast majority of global internet traffic, driving a critical need for efficient audio compression that minimizes network bandwidth without degrading sound quality. Traditional codecs deliver acceptable fidelity at moderate bitrates but degrade significantly at very low bitrates, especially on complex audio such as music. The article evaluates EnCodec, a real-time, high-fidelity neural audio codec designed to compress speech and music across diverse bitrates and sample rates. It aims to demonstrate that deep-learning-based compression can achieve superior acoustic quality over existing industry standards while remaining computationally efficient.
The authors developed an end-to-end convolutional encoder-decoder network combined with residual vector quantization to compress audio into compact discrete representations. To refine training, they introduced a multi-scale spectrogram adversarial loss and a novel gradient balancer mechanism that stabilizes multi-objective optimization by decoupling hyperparameter tuning from loss scales. Optionally, a lightweight Transformer language model was applied for entropy coding to further compress the bitstream. The system was trained on thousands of hours of speech, music, and environmental sounds, and rigorously benchmarked against standard codecs like Opus and EVS as well as recent neural baselines via both objective metrics and subjective human listening evaluations.
The findings establish that EnCodec consistently outperforms standard and neural baselines across all evaluated settings. At 3 kilobits per second, EnCodec achieved higher perceptual quality scores than Opus at 12 kilobits per second and Lyra-v2 at 6 kilobits per second. In stereo music compression at 48 kilohertz, EnCodec at 6 kilobits per second matched the quality of MP3 compression at 64 kilobits per second, achieving a 256-to-1 compression ratio. Additionally, integrating the Transformer language model reduced the required bandwidth by 25% to 40% without perceptual degradation. Computationally, the streaming model operated roughly ten times faster than real time on a single standard computer processing core with an initial algorithmic latency of only 13.3 milliseconds.
These results demonstrate that high-fidelity audio transmission is achievable at unprecedentedly low bitrates, which can substantially reduce infrastructure and distribution costs for music streaming, teleconferencing, and telephony. Operating effectively at very low bitrates makes real-time communication feasible in poor connectivity environments without compromising user experience. For deployment, decision-makers should consider the streamable model for low-latency interactive applications, whereas the non-streamable setup paired with entropy coding provides maximum data savings for offline archiving and on-demand streaming.
Organizations should evaluate pilot implementations of the open-source code for low-bandwidth communication channels and streaming services. Before deploying entropy coding in strict real-time systems, teams must account for modest latency increases and evaluate floating-point precision across varied hardware to avoid decoding discrepancies. Furthermore, while the 24-kilohertz model runs easily in real time on a single CPU, high-resolution 48-kilohertz stereo processing with entropy coding is currently slower than real time on a single CPU core, indicating that hardware acceleration or code optimization is required before production rollout in live high-resolution streaming scenarios.
- Paper: SoundStream: An End-to-End Neural Audio Codec, Neil Zeghidour et al. (2021). SoundStream establishes the core neural audio codec paradigm—combining a fully convolutional encoder-decoder, residual vector quantization, and adversarial/spectral losses—upon which EnCodec directly builds and refines.
- Paper: Neural Discrete Representation Learning, Aäron van den Oord et al. (2017). This foundational paper introduces vector-quantized variational autoencoders (VQ-VAE) and straight-through gradient estimation, providing the underlying discrete representation technique used in neural audio quantization.
- Paper: HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis, Jungil Kong et al. (2020). HiFi-GAN introduces multi-scale and multi-period discriminator formulations along with feature matching losses, serving as a primary foundation for high-fidelity neural audio synthesis and adversarial loss design.
- Paper: Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers, Chengyi Wang et al. (2023). VALL-E directly leverages the discrete acoustic tokens from neural audio codecs like EnCodec to reformulate zero-shot text-to-speech synthesis as a conditional language modeling problem.
- Paper: Qwen3-TTS Technical Report, Hangrui Hu et al. (2026). Qwen3-TTS extends the neural audio tokenization and multi-stage codec generation framework to multilingual, real-time zero-shot voice cloning and dual-track speech synthesis.
