AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Haohe LiuZehua ChenYi YuanXinhao MeiXubo LiuDanilo P. MandicWenwu WangMark D. Plumbley

article2023ICML744 citations

Introduces AudioLDM, a latent diffusion framework that generates high-fidelity audio from text by leveraging contrastive language-audio representations, achieving state-of-the-art generation efficiency on a single GPU while enabling zero-shot text-guided audio manipulation.

Listen

Generating realistic sound effects, music, and speech from natural language descriptions is valuable for virtual reality, gaming, and digital media production. However, existing text-to-audio systems require large-scale, high-quality paired text and audio datasets—which are scarce and often noisy—and demand massive computational resources while yielding limited audio quality.

The article evaluates AudioLDM, a text-to-audio framework designed to generate high-quality audio efficiently by decoupling generative training from cross-modal text-audio alignment. The main objective is to demonstrate that continuous latent diffusion models can be trained entirely on audio-only data while enabling text-guided generation and manipulation during inference.

The authors implemented a continuous latent diffusion model operating within a compressed acoustic space learned by an autoencoder. By leveraging a contrastive language-audio model that shares an aligned embedding space for text and sound, the system trains generative models using audio embeddings alone, bypassing the need for paired descriptive text. The framework was trained across multiple datasets containing up to roughly 3.3 million ten-second audio samples and evaluated against baseline systems using both standardized objective audio distance metrics and blinded evaluations by six audio professionals.

The findings establish that AudioLDM achieves state-of-the-art text-to-audio generation while drastically reducing computational requirements. When evaluated on benchmark data, AudioLDM achieved an objective distance score of 23.31 compared to 47.68 for the primary baseline, DiffSound, indicating roughly double the fidelity and alignment to target data. Subjective human ratings for overall quality and text relevance reached approximately 64 out of 100, markedly outperforming DiffSound's ratings of roughly 44 to 45. In computational efficiency, AudioLDM trained effectively on a single graphics processing unit—compared to 32 to 64 units required by competing models—and synthesized eight ten-second audio clips in under 20 seconds. Furthermore, training on audio-only embeddings outperformed training directly on text-audio pairs, and the system successfully executed zero-shot text-guided style transfer, audio inpainting, and super-resolution without task-specific retraining.

These results demonstrate that developers and organizations can generate and manipulate realistic audio assets with significantly lower infrastructure costs and reduced reliance on curated language captions. Operating in a continuous latent space cuts compute overhead and mitigates risks associated with poor text labeling. However, current limitations include an output sampling rate capped at 16 kHz—which limits musical fidelity—and independent modular training that could cause minor feature misalignments. The authors recommend adopting latent diffusion architectures for commercial sound synthesis and call for future work on end-to-end model optimization, higher sampling rates, and content safeguards to prevent the generation of misleading or harmful audio.

arXiv: 2301.12503
Cover for AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Abstract

Text-to-audio (TTA) system has recently gained attention for its ability to synthesize general audio based on text descriptions. However, previous studies in TTA have limited generation quality with high computational costs. In this study, we propose AudioLDM, a TTA system that is built on a latent space to learn the continuous audio representations from contrastive language-audio pretraining (CLAP) latents. The pretrained CLAP models enable us to train LDMs with audio embedding while providing text embedding as a condition during sampling. By learning the latent representations of audio signals and their compositions without modeling the cross-modal relationship, AudioLDM is advantageous in both generation quality and computational efficiency. Trained on AudioCaps with a single GPU, AudioLDM achieves state-of-the-art TTA performance measured by both objective and subjective metrics (e.g., frechet distance). Moreover, AudioLDM is the first TTA system that enables various text-guided audio manipulations (e.g., style transfer) in a zero-shot fashion. Our implementation and demos are available at this https URL.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Text-Conditional Audio Generation
  • 3.1 Contrastive Language-Audio Pretraining
  • 3.2 Conditional Latent Diffusion Models
  • 3.3 Conditioning Augmentation
  • 3.4 Classifier-free Guidance
  • 3.5 Decoder
  • 4 Text-Guided Audio Manipulation
  • 5 Experiments
  • 5.1 Results
  • 5.2 Ablation Study
  • 6 Conclusions
  • 7 Acknowledgement
  • References
  • A Contrastive Language-Audio Pretraining
  • B Latent Diffusion Model
  • C Variational Autoencoder
  • D Vocoder
  • E Experiment Details
  • F The Effect of Finetuning
  • G Computation Efficiency Comparison
  • H Limitations
  • I Demos

Citation

MLA
Liu, H., et al. “AudioLDM: Text-to-Audio Generation with Latent Diffusion Models”. arXiv, 2023, http://arxiv.org/abs/2301.12503v3.
APA
Liu, H., Chen, Z., Yuan, Y., Mei, X., Liu, X., Mandic, D., Wang, W., & Plumbley, M. D. (2023). AudioLDM: Text-to-Audio Generation with Latent Diffusion Models. arXiv. http://arxiv.org/abs/2301.12503v3
Chicago
Liu, H., Z. Chen, Y. Yuan, et al. 2023. “AudioLDM: Text-to-Audio Generation with Latent Diffusion Models”. arXiv. http://arxiv.org/abs/2301.12503v3.
Harvard
Liu, H. et al. (2023) “AudioLDM: Text-to-Audio Generation with Latent Diffusion Models”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2301.12503v3.
Vancouver
1. Liu H, Chen Z, Yuan Y, Mei X, Liu X, Mandic D, Wang W, Plumbley MD (2023) AudioLDM: Text-to-Audio Generation with Latent Diffusion Models. arXiv

BibTeX

@article{liu2023audioldm,
  title = {AudioLDM: Text-to-Audio Generation with Latent Diffusion Models},
  author = {Liu, Haohe and Chen, Zehua and Yuan, Yi and Mei, Xinhao and Liu, Xubo and Mandic, Danilo and Wang, Wenwu and Plumbley, Mark D.},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2301.12503v3},
  eprint = {2301.12503}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/