AudioLDM: Text-to-Audio Generation with Latent Diffusion Models
Haohe LiuZehua ChenYi YuanXinhao MeiXubo LiuDanilo P. MandicWenwu WangMark D. Plumbley
Introduces AudioLDM, a latent diffusion framework that generates high-fidelity audio from text by leveraging contrastive language-audio representations, achieving state-of-the-art generation efficiency on a single GPU while enabling zero-shot text-guided audio manipulation.
Generating realistic sound effects, music, and speech from natural language descriptions is valuable for virtual reality, gaming, and digital media production. However, existing text-to-audio systems require large-scale, high-quality paired text and audio datasets—which are scarce and often noisy—and demand massive computational resources while yielding limited audio quality.
The article evaluates AudioLDM, a text-to-audio framework designed to generate high-quality audio efficiently by decoupling generative training from cross-modal text-audio alignment. The main objective is to demonstrate that continuous latent diffusion models can be trained entirely on audio-only data while enabling text-guided generation and manipulation during inference.
The authors implemented a continuous latent diffusion model operating within a compressed acoustic space learned by an autoencoder. By leveraging a contrastive language-audio model that shares an aligned embedding space for text and sound, the system trains generative models using audio embeddings alone, bypassing the need for paired descriptive text. The framework was trained across multiple datasets containing up to roughly 3.3 million ten-second audio samples and evaluated against baseline systems using both standardized objective audio distance metrics and blinded evaluations by six audio professionals.
The findings establish that AudioLDM achieves state-of-the-art text-to-audio generation while drastically reducing computational requirements. When evaluated on benchmark data, AudioLDM achieved an objective distance score of 23.31 compared to 47.68 for the primary baseline, DiffSound, indicating roughly double the fidelity and alignment to target data. Subjective human ratings for overall quality and text relevance reached approximately 64 out of 100, markedly outperforming DiffSound's ratings of roughly 44 to 45. In computational efficiency, AudioLDM trained effectively on a single graphics processing unit—compared to 32 to 64 units required by competing models—and synthesized eight ten-second audio clips in under 20 seconds. Furthermore, training on audio-only embeddings outperformed training directly on text-audio pairs, and the system successfully executed zero-shot text-guided style transfer, audio inpainting, and super-resolution without task-specific retraining.
These results demonstrate that developers and organizations can generate and manipulate realistic audio assets with significantly lower infrastructure costs and reduced reliance on curated language captions. Operating in a continuous latent space cuts compute overhead and mitigates risks associated with poor text labeling. However, current limitations include an output sampling rate capped at 16 kHz—which limits musical fidelity—and independent modular training that could cause minor feature misalignments. The authors recommend adopting latent diffusion architectures for commercial sound synthesis and call for future work on end-to-end model optimization, higher sampling rates, and content safeguards to prevent the generation of misleading or harmful audio.
- Paper: High-Resolution Image Synthesis with Latent Diffusion Models, Robin Rombach et al. (2022). This foundational paper introduces Latent Diffusion Models (LDMs), providing the core generative mechanism that AudioLDM adapts for audio representations.
- Paper: DiffWave: A Versatile Diffusion Model for Audio Synthesis, Zhifeng Kong et al. (2021). DiffWave establishes the foundational principles and training formulations for applying diffusion probabilistic models to audio waveform generation.
- Paper: HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis, Jungil Kong et al. (2020). HiFi-GAN introduces high-fidelity neural vocoding essential for converting intermediate continuous spectrogram latents into final raw audio waveforms.
- Paper: Seeing and Hearing: Open-domain Visual-Audio Generation with Diffusion Latent Aligners, Yazhou Xing et al. (2024). This work directly incorporates AudioLDM into a training-free multimodal alignment framework to achieve synchronized audiovisual generation.
- Paper: Fast Timing-Conditioned Latent Audio Diffusion, Zach Evans et al. (2024). This paper advances latent audio diffusion by introducing explicit timing controls and fast long-form stereo generation beyond earlier text-to-audio baselines.
