Fast Timing-Conditioned Latent Audio Diffusion

Zach EvansCJ CarrJosiah TaylorScott H. HawleyJordi Pons

article2024ICML261 citations

Presents Stable Audio, a latent diffusion architecture conditioned on text and timing embeddings to generate variable-length, high-fidelity 44.1kHz stereo music and sound effects of up to 95 seconds in just 8 seconds of inference time.

Listen

Generating high-fidelity, variable-length audio from text descriptions has historically posed severe computational and structural hurdles. Prior generative systems were either constrained to fixed-duration outputs, operated at low audio sampling rates, generated only single-channel mono signals, or suffered from slow generation speeds that hindered practical creative workflows.

The article evaluates Stable Audio, a generative framework designed to rapidly synthesize long-form, variable-length stereo music and sound effects at commercial quality (44.1 kHz) using text prompts and timing controls.

The authors implemented a latent diffusion architecture comprising a fully-convolutional variational autoencoder that compresses stereo audio by a factor of 32, a custom-trained multimodal text encoder, and a 907-million-parameter diffusion network. The model introduces explicit timing embeddings specifying the audio start time and total duration. Performance was benchmarked against leading open-source models using standard public datasets (MusicCaps and AudioCaps), incorporating adapted full-band quantitative metrics alongside human listening evaluations.

The analysis yields several key findings. First, Stable Audio achieves dramatic speed improvements, generating up to 95 seconds of 44.1 kHz stereo audio in just 8 seconds on an enterprise graphics processing unit, outperforming competing autoregressive and diffusion models. Second, the system established superior objective fidelity on music generation benchmarks, scoring 108.69 in statistical audio distance compared to 197.12–354.05 for baseline models. Third, human listeners rated the model highest in audio quality (3.0 out of 4) and text alignment for music. Fourth, unlike alternative models that generate unstructured musical loops, Stable Audio successfully generates coherent musical compositions containing discernible introductions, developments, and outros (achieving 92.1% and 89.4% structural adherence for intros and outros, respectively).

These results demonstrate that commercial-grade, multi-channel audio synthesis can operate in near-real-time without exorbitant computing budgets. By enabling precise duration control and high structural coherence, the technology removes significant technical barriers for creative audio production workflows. However, sound effect text alignment lagged slightly behind dedicated baselines due to dataset imbalances, and spatial correctness on ambient sound effects reached only 57%.

Decision-makers and practitioners deploying this technology should leverage the timing conditioning mechanism to generate custom-length audio assets while applying silence-trimming pipelines for shorter targets. Organizations aiming to improve sound effect generation should expand the diversity and volume of specialized sound effect training data to resolve text alignment gaps. Ongoing governance is also advised to audit potential dataset biases and navigate intellectual property considerations responsibly.

Cover for Fast Timing-Conditioned Latent Audio Diffusion

Abstract

Generating long-form 44.1kHz stereo audio from text prompts can be computationally demanding. Further, most previous works do not tackle that music and sound effects naturally vary in their duration. Our research focuses on the efficient generation of long-form, variable-length stereo music and sounds at 44.1kHz using text prompts with a generative model. Stable Audio is based on latent diffusion, with its latent defined by a fully-convolutional variational autoencoder. It is conditioned on text prompts as well as timing embeddings, allowing for fine control over the content and length of the generated music and sounds. Stable Audio is capable of rendering stereo signals of up to 95 sec at 44.1kHz in 8 sec on an A100 GPU. Despite its compute efficiency and fast inference, it is one of the best in two public text-to-music and -audio benchmarks and, differently from state-of-the-art models, can generate music with structure and stereo sounds.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Architecture
  • 3.1. Variational autoencoder (VAE)
  • 3.2. Conditioning
  • 3.3. Diffusion model
  • 3.4. Inference
  • 4. Training
  • 4.1. Dataset
  • 4.2. Variational autoencoder (VAE)
  • 4.3. Text encoder
  • 4.4. Diffusion model
  • 4.5. Prompt preparation
  • 5. Methodology
  • 5.1. Quantitative metrics
  • 5.2. Qualitative metrics
  • 5.3. Evaluation data
  • 5.4. Baselines
  • 6. Experiments
  • 6.1. How does our autoencoder impact audio fidelity?
  • 6.2. Which text encoder performs the best?
  • 6.3. How accurate is the timing conditioning?
  • 6.4. How does it compare with the state-of-the-art?
  • 6.5. How fast is it?
  • 7. Conclusions
  • Acknowledgments
  • Impact Statement
  • References
  • A. Inference diffusion steps
  • B. MusicCaps and AudioCaps: the original data from Youtube
  • C. Timing conditioning: additional evaluation
  • D. Related work: additional discussion on latent diffusion models
  • E. Additional MusicCaps results: quantitative evaluation without singing-voice prompts
  • F. Implementation details

Knowls

  1. Knowl 1 — Stable Audio Latent Diffusion Framework Overview

    model/method

    Stable Audio is a text-to-audio generative framework designed to synthesize long-form (up to 95.1 seconds), variable-length, full-band stereo music and sound effects at a 44.1 kHz sampling rate. The model executes generative diffusion entirely within the continuous latent space of a pre-trained variational autoencoder (VAE), generating 95 seconds of stereo audio in 8 seconds on an NVIDIA A100 GPU.

    The framework consists of three primary components:

    1. A fully-convolutional stereo Variational Autoencoder (VAE) that downsamples raw stereo waveforms by a temporal factor of 1024 into a 64-channel continuous latent representation, achieving a 32:1 data compression ratio.
    2. A dual conditioning pipeline consisting of (a) penultimate hidden layer embeddings from a custom CLAP text encoder trained with language-audio contrastive loss, and (b) per-second continuous learned timing embeddings encoding the sample's starting second (extseconds_start ext{seconds\_start}) and total source file duration (extseconds_total ext{seconds\_total}).
    3. A 907-million parameter diffusion U-Net operating on the continuous latents. Timestep information modulates network activations via Feature-wise Linear Modulation (FiLM) layers, while text and timing embeddings are concatenated along the sequence dimension and injected through cross-attention layers.
  2. Knowl 2 — Fully-Convolutional Stereo Variational Autoencoder for Audio Compression

    model/method

    The Stable Audio VAE compresses two-channel 44.1 kHz stereo waveforms into continuous latent representations. It employs a fully-convolutional encoder-decoder architecture with 133 million parameters based on the Descript Audio Codec (DAC), omitting vector quantization and using Snake activation functions (f(x)=x+1αsin⁡2(αx)f(x) = x + \frac{1}{\alpha}\sin^2(\alpha x)) to improve waveform reconstruction at high compression factors.

    Given an input audio tensor x∈R2×Lx \in \mathbb{R}^{2 \times L}, where LL is the number of samples, the encoder outputs latents z∈R64×(L/1024)z \in \mathbb{R}^{64 \times (L / 1024)}. The overall data compression ratio is: 2×L64×(L/1024)=32\frac{2 \times L}{64 \times (L / 1024)} = 32

    The VAE is trained on 44.1 kHz stereo audio for 1.1M steps (batch size 256) on 16 NVIDIA A100 GPUs, freezing the encoder after 460k steps and training the decoder for an additional 640k steps. The loss function is: Ltotal=1.0 Lspectral+0.1 Ladv+5.0 Lfm+10−4 LKL\mathcal{L}_{\text{total}} = 1.0 \,\mathcal{L}_{\text{spectral}} + 0.1 \,\mathcal{L}_{\text{adv}} + 5.0 \,\mathcal{L}_{\text{fm}} + 10^{-4} \,\mathcal{L}_{\text{KL}} where:

    • Lspectral\mathcal{L}_{\text{spectral}} is an A-weighted multi-resolution sum and difference STFT loss across window lengths of 2048, 1024, 512, 256, 128, 64, and 32 samples.
    • Ladv\mathcal{L}_{\text{adv}} is an adversarial hinge loss utilizing a multi-scale STFT discriminator modified for stereo inputs across STFT window sizes 2048, 1024, 512, 256, and 128.
    • Lfm\mathcal{L}_{\text{fm}} is the feature matching loss across discriminator layers.
    • LKL\mathcal{L}_{\text{KL}} is the Kullback-Leibler divergence regularizer.
  3. Knowl 3 — Timing Conditioning for Variable-Length Audio Synthesis

    model/method

    Stable Audio enables variable-length generation within a fixed maximum context window (Ltrain=95.1L_{\text{train}} = 95.1 seconds, or 4,194,3044,194,304 samples at 44.1 kHz) by explicitly conditioning on learned timing embeddings.

    During training on an audio file of duration TfileT_{\text{file}} seconds, an extracted chunk starting at offset ToffsetT_{\text{offset}} seconds yields two scalar properties:

    • seconds_start=Toffset\text{seconds\_start} = T_{\text{offset}}: the starting time of the extracted segment within the original recording.
    • seconds_total=Tfile\text{seconds\_total} = T_{\text{file}}: the total duration of the original recording.

    When a training file is shorter than LtrainL_{\text{train}}, seconds_start\text{seconds\_start} is set to 0, seconds_total\text{seconds\_total} is set to the file duration, and the remainder of the 95.195.1-second window is padded with silence.

    Each scalar is converted into a learned continuous per-second embedding vector. The timing embeddings for seconds_start\text{seconds\_start} and seconds_total\text{seconds\_total} are concatenated along the sequence dimension with the text prompt embeddings and passed into the diffusion U-Net's cross-attention layers.

    At inference, requesting an audio of duration T≤LtrainT \le L_{\text{train}} is achieved by setting seconds_start=0\text{seconds\_start} = 0 and seconds_total=T\text{seconds\_total} = T. The model synthesizes TT seconds of coherent audio followed by Ltrain−TL_{\text{train}} - T seconds of trailing silence, which is then removed by trimming.

  4. Knowl 4 — Latent Diffusion U-Net Architecture and Sampling

    model/method

    The latent diffusion model in Stable Audio uses a 907-million parameter U-Net operating on the 64-channel continuous latent space. The network comprises 4 symmetrical levels of downsampling encoder blocks and upsampling decoder blocks connected by residual skip connections, with a 1280-channel bottleneck block:

    • Level 1: 1024 channels, temporal downsampling factor of 1 (no downsampling), 1 self/cross-attention layer per block.
    • Level 2: 1024 channels, downsampling factor of 2, 3 self/cross-attention layers per block.
    • Level 3: 1024 channels, downsampling factor of 2, 3 self/cross-attention layers per block.
    • Level 4: 1280 channels, downsampling factor of 4, 3 self/cross-attention layers per block.
    • Bottleneck: 1280 channels, 3 self/cross-attention layers.

    All attention modules use FlashAttention for memory efficiency over long sequences. Conditioning is applied as follows:

    • Continuous diffusion timesteps modulate activations via FiLM layers.
    • Text and timing tokens are injected via cross-attention layers.

    The U-Net is trained for 640,000 steps with an effective batch size of 256 on 64 NVIDIA A100 GPUs using mixed precision, Exponential Moving Average (EMA), a continuous denoising timestep formulation with a vv-objective, a cosine noise schedule, and 10% conditioning dropout to support classifier-free guidance (CFG).

    Inference uses DPM-Solver++ with 100 diffusion steps and a classifier-free guidance scale of 6.0.

  5. Knowl 5 — In-Domain CLAP Text Conditioning with Penultimate Hidden Layer Features

    model/method

    Stable Audio conditions generation on text prompts using a Contrastive Language-Audio Pretraining (CLAP) model trained directly from scratch on the target training dataset (806,284 audio files and associated metadata). The CLAP architecture consists of:

    • An HTSAT audio encoder with fusion (31M parameters).
    • A RoBERTa text encoder (110M parameters).

    Instead of utilizing pooled embeddings or the final layer's output, conditioning representations are extracted from the penultimate (next-to-last) hidden layer of the RoBERTa text encoder, providing richer feature granularity for cross-attention.

    During training of the CLAP model and diffusion U-Net, metadata (natural language descriptions, BPM, genre, mood, instruments) is formatted into text prompts under two balanced sampling schemes:

    1. Structured format: Field names are included and delimited by vertical bars (e.g., Instruments: Guitar, Drums|Moods: Uplifting).
    2. Unstructured format: Property values are joined by commas without category names (e.g., Guitar, Drums, Uplifting).

    Lists of attributes within metadata types are randomly shuffled to ensure invariance to metadata order.

  6. Knowl 6 — Evaluation Metrics for Long-Form Full-Band Stereo Audio

    model/method

    To evaluate variable-length, long-form full-band (44.1 kHz) stereo audio, three standard metrics were modified:

    1. FDopenl3\text{FD}_{\text{openl3}}: Evaluates sample fidelity and distribution similarity up to 44.1 kHz. Left and right channels of 44.1 kHz stereo audio are separately encoded into OpenL3 embeddings (256 mel bins, 512 embedding dimensions, 0.5 s hop size) and concatenated along the feature dimension to form a combined stereo representation (for mono inputs, mono embeddings are duplicated). The Fréchet distance between generated and ground-truth feature distributions is computed.

    2. KLpasst\text{KL}_{\text{passt}}: Measures semantic label divergence up to 32 kHz. Audios resampled to 32 kHz are divided into overlapping 10-second analysis windows with a 5-second hop (50% overlap). A PaSST audio tagger trained on AudioSet infers class logits for each window. Logits are averaged across all windows prior to softmax normalization, and the Kullback-Leibler divergence between ground-truth and generated label distributions is computed: DKL(Pref∥Pgen)=∑cPref(c)log⁡(Pref(c)Pgen(c))D_{\text{KL}}(P_{\text{ref}} \parallel P_{\text{gen}}) = \sum_{c} P_{\text{ref}}(c) \log \left( \frac{P_{\text{ref}}(c)}{P_{\text{gen}}(c)} \right)

    3. CLAPscore\text{CLAP}_{\text{score}}: Quantifies text-to-audio alignment for durations exceeding 10 seconds. Using the feature-fusion variant of LAION-CLAP at 48 kHz, the audio embedding fuses a 10-second downsampled global representation with three 10-second crops extracted from the start, middle, and end of the audio file. Cosine similarity between prompt text embedding EtextE_{\text{text}} and fused audio embedding EaudioE_{\text{audio}} is computed: CLAPscore=Etext⋅Eaudio∥Etext∥2∥Eaudio∥2\text{CLAP}_{\text{score}} = \frac{E_{\text{text}} \cdot E_{\text{audio}}}{\|E_{\text{text}}\|_2 \|E_{\text{audio}}\|_2}

  7. Knowl 7 — Quantitative Evaluation on MusicCaps and AudioCaps Benchmarks

    data/table

    Stable Audio was evaluated against leading open-source models on the MusicCaps and AudioCaps benchmarks resampled to 44.1 kHz. Models were evaluated on 95-second generation lengths for MusicCaps and 10-second lengths for AudioCaps (for AudioCaps, Stable Audio generated 95-second audio conditioned to 10 seconds and trimmed trailing silence). Inference runtimes reflect a batch size of 1 on a single NVIDIA A100 GPU (40GB VRAM).

    Model Channels / sr Output Length FDopenl3↓\text{FD}_{\text{openl3}} \downarrow KLpasst↓\text{KL}_{\text{passt}} \downarrow CLAPscore↑\text{CLAP}_{\text{score}} \uparrow Inference Time
    MusicCaps
    AudioLDM2-music 1 / 16kHz 95 sec 354.05 1.53 0.30 38 sec
    AudioLDM2-large 1 / 16kHz 95 sec 339.25 1.46 0.30 37 sec
    AudioLDM2-48kHz 1 / 48kHz 95 sec 299.47 2.77 0.22 242 sec
    MusicGen-small 1 / 32kHz 95 sec 205.65 0.96 0.33 126 sec
    MusicGen-large 1 / 32kHz 95 sec 197.12 0.85 0.36 242 sec
    MusicGen-large-stereo 2 / 32kHz 95 sec 216.07 1.04 0.32 295 sec
    Stable Audio 2 / 44.1kHz 95 sec 108.69 0.80 0.46 8 sec
    AudioCaps
    AudioLDM2-large 1 / 16kHz 10 sec 170.31 1.57 0.41 14 sec
    AudioLDM2-48kHz 1 / 48kHz 10 sec 101.11 2.04 0.37 107 sec
    AudioGen-medium 1 / 16kHz 10 sec 186.53 1.42 0.45 36 sec
    Stable Audio 2 / 44.1kHz 95 sec†^\dagger 103.66 2.89 0.24 8 sec

    †^\dagger Generated with 10-second timing conditioning and trailing silence removed.

    Stable Audio achieved the lowest Fréchet Distance (108.69 vs 197.12 for MusicGen-large), lowest label KL divergence (0.80), and highest CLAP alignment score (0.46) on MusicCaps, with an inference time of 8 seconds—substantially faster than the autoregressive baselines (126–295 seconds).

  8. Knowl 8 — Perceptual Evaluation of Audio Quality, Musical Structure, and Stereo Imaging

    data/table

    A blind user study with 19 listeners assessed generated outputs on prompts randomly sampled from MusicCaps and AudioCaps. Audio Quality, Text Alignment, and Musicality were rated on a 5-point Mean Opinion Score (MOS) scale ranging from 0 (Bad) to 4 (Excellent). Stereo Correctness and the presence of compositional structure elements (Intro, Development, Outro) were assessed as binary criteria and reported as percentages.

    MusicCaps AudioCaps
    Metric Stable Audio MusicGen-lg MusicGen-st AudioLDM2 Stable Audio AudioGen AudioLDM2
    Audio Quality 3.0±0.73.0 \pm 0.7 2.1±0.92.1 \pm 0.9 2.8±0.72.8 \pm 0.7 1.2±0.51.2 \pm 0.5 2.5±0.82.5 \pm 0.8 1.3±0.41.3 \pm 0.4 2.2±0.92.2 \pm 0.9
    Text Alignment 2.9±0.82.9 \pm 0.8 2.4±0.92.4 \pm 0.9 2.4±0.92.4 \pm 0.9 1.3±0.61.3 \pm 0.6 2.7±0.92.7 \pm 0.9 2.5±0.92.5 \pm 0.9 2.9±0.82.9 \pm 0.8
    Musicality 2.7±0.92.7 \pm 0.9 2.0±0.92.0 \pm 0.9 2.7±0.92.7 \pm 0.9 1.5±0.71.5 \pm 0.7 — — —
    Stereo correctness 94.7% — 86.8% — 57.0% — —
    Structure: intro 92.1% 36.8% 52.6% 2.6% — — —
    Structure: development 65.7% 68.4% 76.3% 15.7% — — —
    Structure: outro 89.4% 26.3% 15.7% 2.6% — — —

    Stable Audio generated full musical structures with distinct intros (92.1%) and outros (89.4%), whereas baseline models primarily generated ongoing musical development without structured beginnings or endings. Stable Audio also achieved the highest audio quality ratings on both datasets and high stereo correctness (94.7% on music).

  9. Knowl 9 — Text Encoder Ablation Analysis

    empirical result

    An ablation study evaluated the impact of text conditioning representations by training the Stable Audio diffusion model for 350k steps to generate 23-second outputs across three text encoders:

    1. CLAPours\text{CLAP}_{\text{ours}}: CLAP model trained from scratch on the internal dataset (penultimate layer features).
    2. CLAPLAION\text{CLAP}_{\text{LAION}}: Public pre-trained LAION-CLAP text encoder.
    3. T5\text{T5}: Public pre-trained T5 text encoder.

    Quantitative results on 23-second MusicCaps and AudioCaps evaluations:

    MusicCaps (23 sec) AudioCaps (23 sec)
    Conditioning Encoder FDopenl3↓\text{FD}_{\text{openl3}} \downarrow KLpasst↓\text{KL}_{\text{passt}} \downarrow CLAPscore↑\text{CLAP}_{\text{score}} \uparrow FDopenl3↓\text{FD}_{\text{openl3}} \downarrow KLpasst↓\text{KL}_{\text{passt}} \downarrow CLAPscore↑\text{CLAP}_{\text{score}} \uparrow
    Stable Audio w/ CLAPours\text{CLAP}_{\text{ours}} 118.09 0.97 0.44 114.25 2.57 0.16
    Stable Audio w/ CLAPLAION\text{CLAP}_{\text{LAION}} 123.30 1.09 0.43 119.29 2.73 0.19
    Stable Audio w/ T5\text{T5} 126.93 1.06 0.41 119.28 2.69 0.11

    CLAPours\text{CLAP}_{\text{ours}} yielded superior distribution fidelity (lower FDopenl3\text{FD}_{\text{openl3}}) and tag correspondence (lower KLpasst\text{KL}_{\text{passt}}). In-domain contrastive training ensures vocabulary consistency between text representations and the generative latent distribution, eliminating distribution shift.

  10. Knowl 10 — Empirical Precision of Timing Conditioning for Duration Control

    empirical result

    The precision of duration control via timing embeddings was tested by generating audios targeting specific durations (T∈{30,60,90}T \in \{30, 60, 90\} seconds) within the 95.1-second generation window, measuring actual duration using an energy-based silence detector.

    The findings show:

    1. Actual generated audio duration correlates almost perfectly with requested duration across the entire temporal span.
    2. Audio generations exhibit an error of a few seconds with a systematic bias toward finishing slightly before the requested time rather than overshooting it. This ensures that cutting the output signal precisely at the requested duration reliably discards trailing silence without clipping the active audio.
    3. Variance in actual length was slightly elevated in the 40–60 second range due to lower density of training tracks of those specific durations in the dataset.

Coverage note — None was omitted; all contributed architectural components, timing conditioning methods, evaluation metrics, quantitative tables, qualitative human evaluation results, text encoder ablations, and timing accuracy analyses are fully represented.

References

  1. 1.Agostinelli, A., Denk, T. I., Borsos, Z., Engel, J., Verzetti, M., Caillon, A., Huang, Q., Jansen, A., Roberts, A., Tagliasacchi, M., Sharifi, M., Zeghidour, N., and Frank, C. Musiclm: Generating music from text. arXiv, 2023.
  2. 2.Blattmann, A., Rombach, R., Ling, H., Dockhorn, T., Kim, S. W., Fidler, S., and Kreis, K. Align your latents: High-resolution video synthesis with latent diffusion models. arXiv, 2023.
  3. 3.Borsos, Z., Marinier, R., Vincent, D., Kharitonov, E., Pietquin, O., Sharifi, M., Roblek, D., Teboul, O., Grangier, D., Tagliasacchi, M., et al. Audiolm: a language modeling approach to audio generation. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2023.
  4. 4.Chang, H., Zhang, H., Jiang, L., Liu, C., and Freeman, W. T. Maskgit: Masked generative image transformer. CCVPR, 2022.
  5. 5.Copet, J., Kreuk, F., Gat, I., Remez, T., Kant, D., Synnaeve, G., Adi, Y., and Défossez, A. Simple and controllable music generation. arXiv, 2023.
  6. 6.Cramer, A. L., Wu, H.-H., Salamon, J., and Bello, J. P. Look, listen, and learn more: Design choices for deep audio embeddings. ICASSP, 2019.
  7. 7.Dao, T., Fu, D. Y., Ermon, S., Rudra, A., and Ré, C. Flashattention: Fast and memory-efficient exact attention with io-awareness. arXiv, 2022.
  8. 8.Dhariwal, P., Jun, H., Payne, C., Kim, J. W., Radford, A., and Sutskever, I. Jukebox: A generative model for music. arXiv, 2020.
  9. 9.Donahue, C., McAuley, J., and Puckette, M. Adversarial audio synthesis. arXiv, 2018.
  10. 10.Donahue, C., Caillon, A., Roberts, A., Manilow, E., Esling, P., Agostinelli, A., Verzetti, M., Simon, I., Pietquin, O., Zeghidour, N., et al. Singsong: Generating musical accompaniments from singing. arXiv, 2023.
  11. 11.Dong, H.-W., Liu, X., Pons, J., Bhattacharya, G., Pascual, S., Serrà, J., Berg-Kirkpatrick, T., and McAuley, J. Clipsonic: Text-to-audio synthesis with unlabeled videos and pretrained language-vision models. arXiv, 2023.
  12. 12.Défossez, A., Copet, J., Synnaeve, G., and Adi, Y. High fidelity neural audio compression. arXiv, 2022.
  13. 13.Fletcher, H. and Munson, W. A. Loudness, Its Definition, Measurement and Calculation. The Journal of the Acoustical Society of America, 2005.
  14. 14.Forsgren, S. and Martiros, H. Riffusion - stable diffusion for real-time music generation. 2022. URL https://github.com/riffusion/riffusion.
  15. 15.Garcia, H. F., Seetharaman, P., Kumar, R., and Pardo, B. Vampnet: Music generation via masked acoustic token modeling. arXiv, 2023.
  16. 16.Ghosal, D., Majumder, N., Mehrish, A., and Poria, S. Text-to-audio generation using instruction-tuned llm and latent diffusion model. arXiv, 2023.
  17. 17.Hawthorne, C., Simon, I., Roberts, A., Zeghidour, N., Gardner, J., Manilow, E., and Engel, J. Multi-instrument music synthesis with spectrogram diffusion. arXiv, 2022.
  18. 18.Hershey, S., Chaudhuri, S., Ellis, D. P., Gemmeke, J. F., Jansen, A., Moore, R. C., Plakal, M., Platt, D., Saurous, R. A., Seybold, B., et al. Cnn architectures for large-scale audio classification. ICASSP, 2017.
  19. 19.Ho, J., Jain, A., and Abbeel, P. Denoising diffusion probabilistic models. arXiv, 2020.
  20. 20.Huang, Q., Jansen, A., Lee, J., Ganti, R., Li, J. Y., and Ellis, D. P. Mulan: A joint embedding of music audio and natural language. ISMIR, 2022.
  21. 21.Huang, Q., Park, D. S., Wang, T., Denk, T. I., Ly, A., Chen, N., Zhang, Z., Zhang, Z., Yu, J., Frank, C., Engel, J., Le, Q. V., Chan, W., Chen, Z., and Han, W. Noise2music: Text-conditioned music generation with diffusion models. arXiv, 2023a.
  22. 22.Huang, R., Huang, J., Yang, D., Ren, Y., Liu, L., Li, M., Ye, Z., Liu, J., Yin, X., and Zhao, Z. Make-an-audio: Text-to-audio generation with prompt-enhanced diffusion models. arXiv, 2023b.
  23. 23.Kilgour, K., Zuluaga, M., Roblek, D., and Sharifi, M. Fréchet audio distance: A metric for evaluating music enhancement algorithms. arXiv, 2018.
  24. 24.Kim, C. D., Kim, B., Lee, H., and Kim, G. Audiocaps: Generating captions for audios in the wild. Conference of the North American Chapter of the Association for Computational Linguistics, 2019.
  25. 25.Kingma, D. P. and Welling, M. Auto-encoding variational bayes. arXiv, 2013.
  26. 26.Kong, J., Kim, J., and Bae, J. Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis. Advances in Neural Information Processing Systems, 2020.
  27. 27.Koutini, K., Schlüter, J., Eghbal-zadeh, H., and Widmer, G. Efficient training of audio transformers with patchout. Interspeech, 2022.
  28. 28.Kreuk, F., Synnaeve, G., Polyak, A., Singer, U., Défossez, A., Copet, J., Parikh, D., Taigman, Y., and Adi, Y. Audiogen: Textually guided audio generation. arXiv, 2022.
  29. 29.Kumar, R., Seetharaman, P., Luebs, A., Kumar, I., and Kumar, K. High-fidelity audio compression with improved rvqgan. arXiv, 2023.
  30. 30.Lam, M. W., Tian, Q., Li, T., Yin, Z., Feng, S., Tu, M., Ji, Y., Xia, R., Ma, M., Song, X., et al. Efficient neural music generation. NeurIPS, 2024.
  31. 31.Levy, M., Di Giorgi, B., Weers, F., Katharopoulos, A., and Nickson, T. Controllable music production with diffusion models and guidance gradients. arXiv, 2023.
  32. 32.Li, P., Chen, B., Yao, Y., Wang, Y., Wang, A., and Wang, A. Jen-1: Text-guided universal music generation with omnidirectional diffusion models. arXiv, 2023.
  33. 33.Lin, S., Liu, B., Li, J., and Yang, X. Common diffusion noise schedules and sample steps are flawed. IEEE/CVF Winter Conference on Applications of Computer Vision, 2024.
  34. 34.Liu, H., Chen, Z., Yuan, Y., Mei, X., Liu, X., Mandic, D., Wang, W., and Plumbley, M. D. Audioldm: Text-to-audio generation with latent diffusion models. arXiv, 2023a.
  35. 35.Liu, H., Tian, Q., Yuan, Y., Liu, X., Mei, X., Kong, Q., Wang, Y., Wang, W., Wang, Y., and Plumbley, M. D. Audioldm 2: Learning holistic audio generation with self-supervised pretraining. arXiv, 2023b.
  36. 36.Lu, C., Zhou, Y., Bao, F., Chen, J., Li, C., and Zhu, J. Dpm-solver++: Fast solver for guided sampling of diffusion probabilistic models. arXiv, 2022.
  37. 37.Mariani, G., Tallini, I., Postolache, E., Mancusi, M., Cosmo, L., and Rodolà, E. Multi-source diffusion models for simultaneous music generation and separation. arXiv, 2023.
  38. 38.Moliner, E., Lehtinen, J., and Välimäki, V. Solving audio inverse problems with a diffusion model. ICASSP, 2023.
  39. 39.NovelAI. Novelai improvements on stable diffusion, Oct 2022. URL https://shorturl.at/wW034.
  40. 40.Oord, A., Li, Y., Babuschkin, I., Simonyan, K., Vinyals, O., Kavukcuoglu, K., Driessche, G., Lockhart, E., Cobo, L., Stimberg, F., et al. Parallel wavenet: Fast high-fidelity speech synthesis. ICML, 2018.
  41. 41.Oord, A. v. d., Dieleman, S., Zen, H., Simonyan, K., Vinyals, O., Graves, A., Kalchbrenner, N., Senior, A., and Kavukcuoglu, K. Wavenet: A generative model for raw audio. arXiv, 2016.
  42. 42.Parker, J., Spijkervet, J., Kosta, K., Yesiler, F., Kuznetsov, B., Wang, J.-C., Avent, M., Chen, J., and Le, D. Stemgen: A music generation model that listens. ICASSP, 2024.
  43. 43.Pascual, S., Bhattacharya, G., Yeh, C., Pons, J., and Serrà, J. Full-band general audio synthesis with score-based diffusion. ICASSP, 2023.
  44. 44.Pasini, M. and Schlüter, J. Musika! fast infinite waveform music generation. arXiv, 2022.
  45. 45.Perez, E., Strub, F., de Vries, H., Dumoulin, V., and Courville, A. Film: Visual reasoning with a general conditioning layer. arXiv, 2017.
  46. 46.Podell, D., English, Z., Lacey, K., Blattmann, A., Dockhorn, T., Müller, J., Penna, J., and Rombach, R. Sdxl: Improving latent diffusion models for high-resolution image synthesis. arXiv, 2023.
  47. 47.Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I. Learning transferable visual models from natural language supervision. arXiv, 2021.
  48. 48.Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J. Exploring the limits of transfer learning with a unified text-to-text transformer. The Journal of Machine Learning Research, 2020.
  49. 49.Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B. High-resolution image synthesis with latent diffusion models. arXiv, 2022.
  50. 50.Rouard, S. and Hadjeres, G. Crash: Raw audio score-based generative modeling for controllable high-resolution drum sound synthesis. arXiv, 2021.
  51. 51.Salimans, T. and Ho, J. Progressive distillation for fast sampling of diffusion models. arXiv, 2022.
  52. 52.Schneider, F., Jin, Z., and Schölkopf, B. Moûsai: Text-to-music generation with long-context latent diffusion. arXiv, 2023.
  53. 53.Schoeffler, M., Bartoschek, S., Stöter, F.-R., Roess, M., Westphal, S., Edler, B., and Herre, J. webmushra—a comprehensive framework for web-based listening tests. Journal of Open Research Software, 2018.
  54. 54.Sohl-Dickstein, J., Weiss, E. A., Maheswaranathan, N., and Ganguli, S. Deep unsupervised learning using nonequilibrium thermodynamics. arXiv, 2015.
  55. 55.Steinmetz, C. J., Pons, J., Pascual, S., and Serrà, J. Automatic multitrack mixing with a differentiable mixing console of neural audio effects. arXiv, 2020.
  56. 56.Vyas, A., Shi, B., Le, M., Tjandra, A., Wu, Y.-C., Guo, B., Zhang, J., Zhang, X., Adkins, R., Ngan, W., et al. Audiobox: Unified audio generation with natural language prompts. arXiv, 2023.
  57. 57.Wu, Y., Chen, K., Zhang, T., Hui, Y., Berg-Kirkpatrick, T., and Dubnov, S. Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation. ICASSP, 2023.
  58. 58.Yang, D., Tian, J., Tan, X., Huang, R., Liu, S., Chang, X., Shi, J., Zhao, S., Bian, J., Wu, X., et al. Uniaudio: An audio foundation model toward universal audio generation. arXiv, 2023.
  59. 59.Yao, Y., Li, P., Chen, B., and Wang, A. Jen-1 composer: A unified framework for high-fidelity multi-track music generation. arXiv, 2023.
  60. 60.Ziv, A., Gat, I., Lan, G. L., Remez, T., Kreuk, F., Défossez, A., Copet, J., Synnaeve, G., and Adi, Y. Masked audio generation using a single non-autoregressive transformer. arXiv, 2024.
  61. 61.Ziyin, L., Hartwig, T., and Ueda, M. Neural networks fail to learn periodic functions and how to fix it. arXiv, 2020.

Citation

MLA
Evans, Z., et al. “Fast Timing-Conditioned Latent Audio Diffusion”. arXiv, 2024, http://arxiv.org/abs/2402.04825v3.
APA
Evans, Z., Carr, C., Taylor, J., Hawley, S. H., & Pons, J. (2024). Fast Timing-Conditioned Latent Audio Diffusion. arXiv. http://arxiv.org/abs/2402.04825v3
Chicago
Evans, Z., C. Carr, J. Taylor, S. H. Hawley, and J. Pons. 2024. “Fast Timing-Conditioned Latent Audio Diffusion”. arXiv. http://arxiv.org/abs/2402.04825v3.
Harvard
Evans, Z. et al. (2024) “Fast Timing-Conditioned Latent Audio Diffusion”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2402.04825v3.
Vancouver
1. Evans Z, Carr C, Taylor J, Hawley SH, Pons J (2024) Fast Timing-Conditioned Latent Audio Diffusion. arXiv

BibTeX

@article{evans2024fast,
  title = {Fast Timing-Conditioned Latent Audio Diffusion},
  author = {Evans, Zach and Carr, CJ and Taylor, Josiah and Hawley, Scott H. and Pons, Jordi},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2402.04825v3},
  eprint = {2402.04825}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/