Symbolic Music Generation with Non-Differentiable Rule Guided Diffusion

Yujia HuangAdishree GhatareYuanzhe LiuZiniu HuQinsheng ZhangChandramouli Shama SastrySiddharth GururaniSageev OoreYisong Yue

article2024ICML47 citations

Introduces Stochastic Control Guidance, a training-free method that enables pre-trained diffusion models to adhere to non-differentiable symbolic musical rules using only forward function evaluations.

Listen

Generating symbolic music (such as piano rolls) with modern generative models offers significant potential for human-AI co-creation. However, giving human composers precise control over specific musical rules—such as target chord progressions or note densities—remains a fundamental challenge. Existing diffusion-based models require differentiable loss functions to calculate gradients for guidance or necessitate training task-specific surrogate models, which are resource-intensive, difficult to train accurately for complex rules, and impractical for flexible, multi-rule compositions.

The article introduces and evaluates Stochastic Control Guidance (SCG), a training-free, plug-and-play guidance framework that controls pre-trained diffusion models using non-differentiable and black-box rules. It pairs this method with a high-resolution latent diffusion architecture designed to capture expressive piano performances with 10-millisecond timing precision.

To achieve non-differentiable control, the authors formulate the generation process as a stochastic optimal control problem. Rather than calculating mathematical gradients, SCG generates multiple candidate realizations at each intermediate generation step, estimates the final clean output for each candidate, and selects the direction that minimizes rule violation based purely on forward evaluation. The system was trained on approximately 1,000 hours of diverse classical and pop piano datasets and systematically evaluated across unconditional generation, individual rule guidance, composite multi-rule guidance, and music editing tasks against competitive baselines.

The findings show that SCG provides superior controllability over non-differentiable musical rules. In individual rule testing on non-differentiable note density, SCG lowered the error metric from 2.486 (unconditional) down to 0.131, markedly outperforming surrogate classifier methods (0.698) and neural-network-based loss guidance (1.261). For chord progression adherence, SCG reduced the error metric to 0.273, compared to 0.723 for classifier guidance. In composite multi-rule settings, combining SCG with lightweight classifier guidance produced the best overall rule adherence across all target rules simultaneously while maintaining high musical fidelity. Additionally, human listening surveys with 15 experienced musicians ranked the proposed system highest across rule alignment, musical creativity, coherence, and overall aesthetic appeal.

These results demonstrate that generative models can achieve fine-grained symbolic control without retraining or relying on fragile differentiable surrogates. In practice, this eliminates the high compute costs associated with maintaining custom models for every desired musical rule. It also enables broader applications across fields where non-differentiable constraints or black-box physics simulators must guide generative processes.

For practical implementation, practitioners and creators can adopt SCG as a plug-and-play compositional tool. When low latency is required, teams should combine coarse classifier guidance with a modest candidate sample size (e.g., four samples per step) or apply stochastic fast-sampling algorithms, which achieve comparable control while speeding up generation roughly fourfold.

The primary limitation of SCG is its increased computational cost during sampling due to generating and evaluating multiple candidates per step. Users must also manage an inherent trade-off where excessively optimizing a single constraint can degrade overall musical coherence. Nevertheless, the underlying theoretical formulation and extensive empirical validations provide strong confidence in SCG as an effective method for non-differentiable controlled generation.

arXiv: 2402.14285yjhuangcd/rule-guided-music

No sufficiently relevant recommendations were found.

Cover for Symbolic Music Generation with Non-Differentiable Rule Guided Diffusion

Abstract

We study the problem of symbolic music generation (e.g., generating piano rolls), with a technical focus on non-differentiable rule guidance. Musical rules are often expressed in symbolic form on note characteristics, such as note density or chord progression, many of which are non-differentiable which pose a challenge when using them for guided diffusion. We propose Stochastic Control Guidance (SCG), a novel guidance method that only requires forward evaluation of rule functions that can work with pre-trained diffusion models in a plug-and-play way, thus achieving training-free guidance for non-differentiable rules for the first time. Additionally, we introduce a latent diffusion architecture for symbolic music generation with high time resolution, which can be composed with SCG in a plug-and-play fashion. Compared to standard strong baselines in symbolic music generation, this framework demonstrates marked advancements in music quality and rule-based controllability, outperforming current state-of-the-art generators in a variety of settings. For detailed demonstrations, code and model checkpoints, please visit our project website.

Table of Contents

  • 1. Introduction
  • 2. Related Works
  • 3. Background
  • 4. Non-Differentiable Rule Guidance
  • 4.1. Rule Guidance Problem
  • 4.2. Guidance via Stochastic Control
  • 4.3. Practical Algorithms
  • 4.4. General Theoretical Connection
  • 5. Latent Diffusion Architecture
  • 6. Experiments
  • 6.1. Experimental Settings
  • 6.2. Unconditional Generation
  • 6.3. Individual Rule Guidance
  • 6.4. Composite Rule Guidance
  • 6.5. Ablation Studies
  • 6.6. Subjective Evaluation
  • 6.7. Examples of Our System as a Compositional Tool
  • 7. Conclusion and Future directions
  • Acknowledgements
  • Impact Statement
  • References
  • A. Proofs
  • A.1. Proof of Theorem 4.1
  • A.2. Proof of Proposition 4.2
  • B. Compatibility of SCG with Various Sampling Procedures
  • C. Additional Experiment Results
  • C.1. Unconditional Generation
  • C.2. Editing
  • D. Detailed Experiment Setup
  • D.1. Music Rules
  • D.2. Training Setup
  • D.3. Objective Evaluation Setup
  • E. Training Surrogate Models for Music Rules
  • F. Losses over Stochastic Control Guided Sampling Process
  • G. More Ablation Studies
  • G.1. Unconditional Generation
  • G.2. Latent vs Pixel Space
  • G.3. Composite Rule Guidance
  • H. Rule-Guided Generation Survey
  • H.1. Details
  • H.2. Survey

Knowls

  1. Knowl 1 — Stochastic Control Guidance sampling

    algorithm

    Stochastic Control Guidance (SCG) steers a pretrained stochastic diffusion sampler toward samples with low rule loss without differentiating through the rule. Its inputs are a pretrained DDPM noise predictor ϵθ\epsilon_\theta, a rule target yy, a forward-evaluable loss ℓy\ell_y, and a candidate count nn; its output is a generated sample. In the reported default setup, DDPM uses 1000 reverse steps, SCG begins at timestep t=750t=750, and n=16n=16. The state xtx_t is the noisy diffusion sample at timestep tt, βt\beta_t is the forward variance, αt=1−βt\alpha_t=1-\beta_t, αˉt=∏s=1tαs\bar\alpha_t=\prod_{s=1}^t\alpha_s, and σt\sigma_t is the sampler's DDPM transition standard deviation. For latent diffusion, DD denotes the VAE decoder; for pixel-space diffusion, use the sample directly when evaluating the rule.

    Input: Pretrained DDPM noise predictor epsilon_theta, rule loss ell_y, target y, candidate count n = 16, T = 1000
    Initialize x_T from a standard Gaussian
    For t = T down to 1:
        Compute DDPM posterior mean m_t = (x_t - beta_t epsilon_theta(x_t, t) / sqrt(1 - alpha_bar_t)) / sqrt(alpha_t)
        If t <= 750 and t > 1:
            For i = 1 to n:
                Sample z_i from a standard Gaussian
                Set x_i_(t-1) = m_t + sigma_t z_i
                Estimate clean candidate xhat_i_0 = (x_i_(t-1) - sqrt(1 - alpha_bar_(t-1)) epsilon_theta(x_i_(t-1), t-1)) / sqrt(alpha_bar_(t-1))
                Evaluate score_i = -ell_y(D(xhat_i_0), y), using D as the identity for pixel-space samples
            Select k = argmax_i score_i
            Set x_(t-1) = x_k_(t-1)
        Else:
            Take an ordinary DDPM transition from x_t to x_(t-1); at t = 1 use m_t
    Return D(x_0) for latent diffusion, or x_0 for pixel-space diffusion

    At each guided step, SCG chooses the candidate whose estimated clean sample best satisfies the rule. It needs only forward evaluations of ℓy\ell_y, so it accommodates non-differentiable rules and black-box rule functions.

  2. Knowl 2 — Optimal-control interpretation of rule guidance

    theoretical result

    A reverse diffusion sampler can be viewed as an uncontrolled stochastic process that is modified by a control to favor samples satisfying a rule. Let ηt\eta_t be the time-reparameterized diffusion state, with initial state η0∼N(0,I)\eta_0\sim\mathcal{N}(0,I) and uncontrolled dynamics dηt=f~(ηt,t) dt+g(t) dwtd\eta_t=\tilde f(\eta_t,t)\,dt+g(t)\,dw_t. Here f~\tilde f is the uncontrolled drift, g(t)g(t) is the scalar diffusion coefficient, and wtw_t is standard Wiener noise. A control utu_t changes the dynamics to dηt=f~ dt+g(t)(ut dt+dwt)d\eta_t=\tilde f\,dt+g(t)(u_t\,dt+dw_t) and incurs the expected cost of terminal rule loss plus quadratic control effort, E[ℓy(ηT)+12∫0T∥ut∥2dt]\mathbb{E}[\ell_y(\eta_T)+\tfrac12\int_0^T\|u_t\|^2dt]. The rule likelihood is assumed to satisfy p(y∣ηT)∝e−ℓy(ηT)p(y\mid\eta_T)\propto e^{-\ell_y(\eta_T)}.

    Define the desirability function under the uncontrolled process Q0Q^0 as Ψ(η,t)=EQ0[e−ℓy(ηT)∣ηt=η]\Psi(\eta,t)=\mathbb{E}_{Q^0}[e^{-\ell_y(\eta_T)}\mid\eta_t=\eta]. The optimal control is ut∗(η)=g(t)∇ηlog⁡Ψ(η,t)u_t^*(\eta)=g(t)\nabla_\eta\log\Psi(\eta,t). The paper establishes that this control yields terminal samples from the desired conditional distribution, Q∗(ηT)=p(ηT∣y)Q^*(\eta_T)=p(\eta_T\mid y), and that Ψ(ηt,t)=c p(y∣ηt)\Psi(\eta_t,t)=c\,p(y\mid\eta_t) for a constant cc independent of ηt\eta_t. Thus, classifier guidance and differentiable loss guidance can be interpreted as ways of estimating or approximating the same likelihood-based optimal control; SCG instead uses forward evaluations of the rule.

  3. Knowl 3 — One-step approximation used by SCG

    theoretical result

    Computing the exact path-integral optimal control would require accounting for full future trajectories from each possible diffusion update. SCG replaces that costly calculation with a one-step choice: for a candidate update to ηt+dt\eta_{t+dt}, it estimates the terminal clean state by the posterior mean η^T=E[ηT∣ηt+dt]\hat\eta_T=\mathbb{E}[\eta_T\mid\eta_{t+dt}], obtained using a Tweedie-style denoising estimate, and selects the update maximizing −ℓy(η^T)-\ell_y(\hat\eta_T). Here ηt\eta_t is the current state, dtdt is an infinitesimal step, and ηT\eta_T is the terminal data-space state. In the paper's limiting interpretation, scaling terminal cost as ℓy/K\ell_y/K and taking K→0K\to0 makes the full path objective favor trajectories with low terminal loss; selecting by the posterior-mean estimate is presented as optimizing a lower bound on that objective. This approximation avoids unrolling future trajectories and does not require a gradient of the rule loss.

  4. Knowl 4 — High-resolution latent diffusion architecture for piano music

    model/method

    The symbolic music generator represents each 10 ms time column with three channels: piano-roll note velocity in the range 0–127, a binary onset indicator, and sustain-pedal control. A VAE encodes chunks of shape 3×128×1283\times128\times128 into latent chunks of shape 4×16×164\times16\times16. Concatenating eight consecutive chunks gives a latent representation of shape 4×16×1284\times16\times128, corresponding to a 10.24-second excerpt. The VAE uses an L1L_1 reconstruction objective, KL regularization, and a denoising reconstruction objective trained to undo musical perturbations such as note shifts, added adjacent notes, timing shifts, and note deletion; perturbations affect at most 25% of notes. The VAE is trained for 800,000 steps.

    A DiT-XL transformer diffusion model is trained on the concatenated latent representation, reshaped into a sequence of 256 tokens with 32 dimensions before projection to the transformer's hidden dimension of 1152. Rotary positional embeddings support variable sequence lengths, and DiffCollage combines score functions for shorter segments to generate longer excerpts. The diffusion model is trained for 1.2 million steps with dataset conditioning for classical performance, classical sheet music, and pop piano; classifier-free condition dropout is 0.1. The training data comprise MAESTRO (about 200 hours), roughly 14,000 Muscore MIDI files (about 700 hours), Pop1k7 (108 hours), and Pop909 (60 hours).

  5. Knowl 5 — Operational definitions of the guided music rules

    definition

    The evaluated rules are defined on piano-roll excerpts with 10 ms time columns. Pitch histogram (PH) is a velocity-weighted histogram over the 12 pitch classes across the excerpt; its target is a 12-dimensional vector. Note density (ND) is measured in 128×128128\times128 windows, each spanning 1.28 seconds. Its target is a 16-dimensional vector: eight vertical-density values and eight horizontal-density values. For a window of T=128T=128 time columns, vertical density is NDvertical=1T∑t=1Tnon(t)\mathrm{ND}_{\mathrm{vertical}}=\frac{1}{T}\sum_{t=1}^{T}n_{\mathrm{on}}(t), where non(t)n_{\mathrm{on}}(t) is the number of sounding notes at time tt; horizontal density is NDhorizontal=∑t=1T1(nstart(t)≥1)\mathrm{ND}_{\mathrm{horizontal}}=\sum_{t=1}^{T}\mathbf{1}(n_{\mathrm{start}}(t)\ge1), where nstart(t)n_{\mathrm{start}}(t) counts note onsets at time tt. Chord progression (CP) is extracted with the music21 chord-analysis tool and grouped into seven classes. The target specifies eight chords, one per 128×128128\times128 window, selected as the longest chord in that window.

  6. Knowl 6 — SCG improves guidance for non-differentiable music rules

    empirical result

    The individual-rule evaluation used 200 Muscore test excerpts as targets and generated 200 conditioned excerpts per rule; SCG used 16 candidates per guided step. Lower loss indicates closer rule matching, while higher overlapping area (OA) indicates a closer match between generated and rule-matched music distributions. PH denotes pitch histogram, ND note density, and CP chord progression. Values are mean ±\pm standard deviation.

    Method PH loss ↓\downarrow PH OA ↑\uparrow ND loss ↓\downarrow ND OA ↑\uparrow CP loss ↓\downarrow CP OA ↑\uparrow
    No Guidance 0.018 ±\pm 0.010 0.842 ±\pm 0.012 2.486 ±\pm 3.530 0.830 ±\pm 0.016 0.831 ±\pm 0.142 0.854 ±\pm 0.026
    Classifier 0.005 ±\pm 0.004 0.855 ±\pm 0.020 0.698 ±\pm 0.587 0.861 ±\pm 0.025 0.723 ±\pm 0.200 0.850 ±\pm 0.033
    DPS-NN 0.001 ±\pm 0.002 0.849 ±\pm 0.018 1.261 ±\pm 2.340 0.667 ±\pm 0.113 0.414 ±\pm 0.256 0.839 ±\pm 0.039
    DPS-Rule 0.010 ±\pm 0.008 0.635 ±\pm 0.006 2.508 ±\pm 2.798 0.800 ±\pm 0.080 – –
    SCG 0.003 ±\pm 0.004 0.867 ±\pm 0.005 0.131 ±\pm 0.325 0.842 ±\pm 0.031 0.273 ±\pm 0.1637 0.851 ±\pm 0.011

    SCG achieves the lowest reported loss for the two non-differentiable rules, ND and CP, without a trained surrogate model. For PH, which is differentiable, DPS-NN has lower loss (0.001 versus 0.003). SCG's OA is not uniformly best, illustrating that stronger control of a specified rule need not improve every measure of musical-distribution similarity.

  7. Knowl 7 — Composite rule guidance and hybrid control

    empirical result

    Composite guidance was evaluated by conditioning on PH, ND, and CP targets extracted from 200 Muscore test excerpts. The SCG objective weighted the three losses by 40, 1, and 1, respectively; classifier-based methods combined per-rule guidance signals. The table reports rule losses (lower is better) and OA (higher is better), as mean ±\pm standard deviation. Retrieval selects existing dataset examples and serves as a non-generative reference; the other named baselines are conditional music generators or diffusion guidance methods.

    Method PH loss ↓\downarrow ND loss ↓\downarrow CP loss ↓\downarrow OA ↑\uparrow
    Retrieval 0.006 ±\pm 0.005 0.433 ±\pm 1.068 0.556 ±\pm 0.182 0.886 ±\pm 0.005
    Figaro expert 0.007 ±\pm 0.007 2.303 ±\pm 2.256 0.761 ±\pm 0.187 0.857 ±\pm 0.047
    Figaro expert+learned 0.006 ±\pm 0.009 1.489 ±\pm 2.737 0.726 ±\pm 0.214 0.883 ±\pm 0.008
    MuseCoCo 0.040 ±\pm 0.026 2.734 ±\pm 3.551 0.821 ±\pm 0.163 0.753 ±\pm 0.038
    No Guidance 0.018 ±\pm 0.010 2.486 ±\pm 3.530 0.831 ±\pm 0.142 0.803 ±\pm 0.096
    Classifier 0.006 ±\pm 0.006 0.822 ±\pm 0.844 0.724 ±\pm 0.205 0.859 ±\pm 0.026
    DPS-NN 0.004 ±\pm 0.006 1.366 ±\pm 2.265 0.661 ±\pm 0.257 0.752 ±\pm 0.079
    SCG 0.014 ±\pm 0.009 0.466 ±\pm 0.648 0.446 ±\pm 0.205 0.874 ±\pm 0.007
    DPS-NN + SCG 0.002 ±\pm 0.007 0.238 ±\pm 0.531 0.313 ±\pm 0.231 0.781 ±\pm 0.084
    Classifier + SCG 0.003 ±\pm 0.005 0.148 ±\pm 0.203 0.284 ±\pm 0.197 0.864 ±\pm 0.010

    Among the diffusion guidance methods, combining classifier guidance with SCG gives the lowest losses for all three rules in this comparison, with OA 0.864. The hybrid uses a learned gradient signal to provide a preliminary direction and SCG to select among stochastic candidates, improving the rule losses over classifier guidance alone. Satisfying multiple constraints together is harder than guiding a single rule under the same sampling budget.

  8. Knowl 8 — Unconditional generation across musical datasets

    empirical result

    Unconditional-generation quality was assessed by average overlapping area (OA) over seven musical attributes, comparing generated samples with held-out data from the corresponding dataset; larger OA indicates closer attribute-distribution overlap. Each method generated 400 excerpts of 10.24 seconds, and the reported values are mean ±\pm standard deviation. GT is a comparison between two subsets of real data, not a generative model. SCG's latent diffusion model performs strongly across classical performance (Maestro), classical sheet music (Muscore), and pop music (Pop), with the highest OA among the listed generators on all three datasets.

    Dataset GT MusicTr Remi CPW PolyDiff SCG
    Maestro 0.944 ±\pm 0.002 0.903 ±\pm 0.005 0.847 ±\pm 0.005 0.801 ±\pm 0.006 0.842 ±\pm 0.007 0.943 ±\pm 0.003
    Muscore 0.945 ±\pm 0.004 0.901 ±\pm 0.004 0.879 ±\pm 0.006 0.843 ±\pm 0.007 0.845 ±\pm 0.004 0.934 ±\pm 0.003
    Pop 0.957 ±\pm 0.002 0.845 ±\pm 0.004 0.866 ±\pm 0.004 0.899 ±\pm 0.005 0.883 ±\pm 0.004 0.939 ±\pm 0.004

    SCG is near the real-data reference on Maestro and achieves higher OA than the listed generators on Muscore and Pop. The different baseline rankings across classical and pop datasets contrast with SCG's consistent performance across the three data distributions.

  9. Knowl 9 — Candidate count trades sampling time against control and quality

    empirical result

    This ablation guides samples toward target note density and measures rule loss, OA against the full test dataset (OA full), OA against rule-matched data (OA), and elapsed time to generate a batch of four excerpts. Values are mean ±\pm standard deviation where reported. Increasing SCG's candidate count reduces note-density loss but costs more sampling time; combining classifier guidance with SCG improves the trade-off at small candidate counts.

    Method Candidates per step Loss ↓\downarrow OA full ↑\uparrow OA ↑\uparrow Time (s)
    No Guidance – 2.486 ±\pm 3.530 0.918 ±\pm 0.005 0.830 ±\pm 0.016 25.4
    Source – 0 0.923 ±\pm 0.008 – –
    Classifier 1 0.698 ±\pm 0.587 0.914 ±\pm 0.006 0.861 ±\pm 0.025 47.8
    DPS-NN 1 1.261 ±\pm 2.340 0.735 ±\pm 0.012 0.667 ±\pm 0.113 109.3
    SCG 4 0.318 ±\pm 0.770 0.895 ±\pm 0.006 0.873 ±\pm 0.023 277.7
    SCG 8 0.214 ±\pm 0.368 0.877 ±\pm 0.006 0.847 ±\pm 0.014 531.6
    SCG 16 0.131 ±\pm 0.325 0.880 ±\pm 0.003 0.842 ±\pm 0.031 1242.6
    Classifier + SCG 4 0.151 ±\pm 0.298 0.906 ±\pm 0.006 0.861 ±\pm 0.011 301.9
    Classifier + SCG 8 0.098 ±\pm 0.179 0.893 ±\pm 0.004 0.839 ±\pm 0.024 555.6
    Classifier + SCG 16 0.064 ±\pm 0.159 0.899 ±\pm 0.007 0.849 ±\pm 0.018 1253.9

    SCG with 16 candidates lowers loss relative to SCG with 4 candidates but takes 1242.6 seconds rather than 277.7 seconds for four outputs. Classifier + SCG with four candidates reaches loss 0.151 in 301.9 seconds, close to SCG alone with 16 candidates (loss 0.131) at roughly one quarter of the time. The OA measures do not improve monotonically with candidate count, consistent with a trade-off between precise rule matching and broader musical-distribution similarity.

  10. Knowl 10 — Listening evaluation of rule-guided samples

    empirical result

    A listening study compared SCG with classifier guidance and DPS using 12 generated piano excerpts, each 10.24 seconds long. The excerpts covered four rule sets, each comprising pitch histogram, note density, and chord progression constraints. Fifteen participants with substantial musical experience rated rule alignment, creativity, coherence, and overall preference. SCG received higher average scores than the two baselines in each of the four evaluation dimensions. The study therefore provides human-rating evidence that stronger rule control was accompanied by favorable judgments of musicality in this evaluation; it does not establish that every individual output or listener preferred SCG.

Coverage note — The musician-led improvisation demonstrations and detailed editing benchmark are omitted as secondary application evidence; the core method, architecture, rule-control evaluations, and principal quality trade-offs are represented.

References

  1. 1.Atassi, L. Generating symbolic music using diffusion models. arXiv preprint arXiv:2303.08385, 2023.
  2. 2.Berner, J., Richter, L., and Ullrich, K. An optimal control perspective on diffusion-based generative modeling. arXiv preprint arXiv:2211.01364, 2022.
  3. 3.Brunner, G., Konrad, A., Wang, Y., and Wattenhofer, R. Midi-vae: Modeling dynamics and instrumentation of music with applications to style transfer. 19th International Society for Music Information Retrieval Conference, 2018.
  4. 4.Choi, J., Kim, S., Jeong, Y., Gwon, Y., and Yoon, S. Ilvr: Conditioning method for denoising diffusion probabilistic models. ICCV, 2021.
  5. 5.Choi, K., Hawthorne, C., Simon, I., Dinculescu, M., and Engel, J. Encoding musical style with transformer autoencoders. In International Conference on Machine Learning, pp. 1899–1908. PMLR, 2020.
  6. 6.Chung, H., Kim, J., Mccann, M. T., Klasky, M. L., and Ye, J. C. Diffusion posterior sampling for general noisy inverse problems. ICLR, 2023.
  7. 7.Cuthbert, M. S. and Ariza, C. music21: A toolkit for computer-aided musicology and symbolic music data. 2010.
  8. 8.Dai Pra, P. A stochastic control approach to reciprocal diffusion processes. Applied mathematics and Optimization, 23:313–329, 1991.
  9. 9.Dhariwal, P. and Nichol, A. Diffusion models beat gans on image synthesis. Advances in neural information processing systems, 34:8780–8794, 2021.
  10. 10.Dong, H.-W., Hsiao, W.-Y., Yang, L.-C., and Yang, Y.-H. Musegan: Multi-track sequential generative adversarial networks for symbolic music generation and accompaniment. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 32, 2018.
  11. 11.Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al. An image is worth 16x16 words: Transformers for image recognition at scale. ICLR, 2021.
  12. 12.Efron, B. Tweedie’s formula and selection bias. Journal of the American Statistical Association, 106(496):1602–1614, 2011.
  13. 13.Evans, L. C. Partial differential equations, volume 19. American Mathematical Society, 2022.
  14. 14.Fleming, W. H. and Mitter, S. K. Optimal control and nonlinear filtering for nondegenerate diffusion processes. Stochastics: An International Journal of Probability and Stochastic Processes, 8(1):63–77, 1982.
  15. 15.Hawthorne, C., Stasyuk, A., Roberts, A., Simon, I., Huang, C.-Z. A., Dieleman, S., Elsen, E., Engel, J., and Eck, D. Enabling factorized piano music modeling and generation with the MAESTRO dataset. In International Conference on Learning Representations, 2019.
  16. 16.Ho, J. and Salimans, T. Classifier-free diffusion guidance. arXiv preprint arXiv:2207.12598, 2022.
  17. 17.Ho, J., Jain, A., and Abbeel, P. Denoising diffusion probabilistic models. Advances in neural information processing systems, 33:6840–6851, 2020.
  18. 18.Ho, J., Chan, W., Saharia, C., Whang, J., Gao, R., Gritsenko, A., Kingma, D. P., Poole, B., Norouzi, M., Fleet, D. J., et al. Imagen video: High definition video generation with diffusion models. arXiv preprint arXiv:2210.02303, 2022.
  19. 19.Hsiao, W.-Y., Liu, J.-Y., Yeh, Y.-C., and Yang, Y.-H. Compound word transformer: Learning to compose full-song music over dynamic directed hypergraphs. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, pp. 178–186, 2021.
  20. 20.Huang, C.-Z. A., Vaswani, A., Uszkoreit, J., Shazeer, N., Simon, I., Hawthorne, C., Dai, A. M., Hoffman, M. D., Dinculescu, M., and Eck, D. Music transformer. arXiv preprint arXiv:1809.04281, 2018.
  21. 21.Huang, Q., Park, D. S., Wang, T., Denk, T. I., Ly, A., Chen, N., Zhang, Z., Zhang, Z., Yu, J., Frank, C., et al. Noise2music: Text-conditioned music generation with diffusion models. arXiv preprint arXiv:2302.03917, 2023.
  22. 22.Huang, Y.-S. and Yang, Y.-H. Pop music transformer: Beat-based modeling and generation of expressive pop piano compositions. In Proceedings of the 28th ACM international conference on multimedia, pp. 1180–1188, 2020.
  23. 23.Kappen, H. J. Path integrals and symmetry breaking for optimal control theory. Journal of statistical mechanics: theory and experiment, 2005(11):P11011, 2005.
  24. 24.Kawar, B., Elad, M., Ermon, S., and Song, J. Denoising diffusion restoration models. Advances in Neural Information Processing Systems, 35:23593–23606, 2022.
  25. 25.Li, S. and Sung, Y. Melodydiffusion: Chord-conditioned melody generation using a transformer-based diffusion model. Mathematics, 11(8):1915, 2023.
  26. 26.Loshchilov, I. and Hutter, F. Decoupled weight decay regularization. ICLR, 2019.
  27. 27.Lu, P., Xu, X., Kang, C., Yu, B., Xing, C., Tan, X., and Bian, J. Musecoco: Generating symbolic music from text. arXiv preprint arXiv:2306.00110, 2023.
  28. 28.Meade, N., Barreyre, N., Lowe, S. C., and Oore, S. Exploring conditioning for generative music systems with human-interpretable controls. arXiv preprint arXiv:1907.04352, 2019.
  29. 29.Meng, C., Song, Y., Song, J., Wu, J., Zhu, J.-Y., and Ermon, S. SDEdit: Image synthesis and editing with stochastic differential equations. ICLR, 2022.
  30. 30.Min, L., Jiang, J., Xia, G., and Zhao, J. Polyffusion: A diffusion model for polyphonic score generation with internal and external controls. Proc. of the 24th Int. Society for Music Information Retrieval Conf, 2023.
  31. 31.Øksendal, B. Stochastic differential equations. Springer, 2003.
  32. 32.Pavon, M. Stochastic control and nonequilibrium thermodynamical systems. Applied Mathematics and Optimization, 19:187–202, 1989.
  33. 33.Peebles, W. and Xie, S. Scalable diffusion models with transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 4195–4205, 2023.
  34. 34.Ren, Y., He, J., Tan, X., Qin, T., Zhao, Z., and Liu, T.-Y. Popmag: Pop music accompaniment generation. In Proceedings of the 28th ACM international conference on multimedia, pp. 1198–1206, 2020.
  35. 35.Roberts, A., Engel, J., Raffel, C., Hawthorne, C., and Eck, D. A hierarchical latent vector model for learning long-term structure in music. In International conference on machine learning, pp. 4364–4373. PMLR, 2018.
  36. 36.Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 10684–10695, 2022.
  37. 37.Song, J., Meng, C., and Ermon, S. Denoising diffusion implicit models. ICLR, 2021a.
  38. 38.Song, J., Zhang, Q., Yin, H., Mardani, M., Liu, M.-Y., Kautz, J., Chen, Y., and Vahdat, A. Loss-guided diffusion models for plug-and-play controllable generation. In International Conference on Machine Learning, 2023.
  39. 39.Song, Y., Sohl-Dickstein, J., Kingma, D. P., Kumar, A., Ermon, S., and Poole, B. Score-based generative modeling through stochastic differential equations. ICLR, 2021b.
  40. 40.Su, J., Ahmed, M., Lu, Y., Pan, S., Bo, W., and Liu, Y. Roformer: Enhanced transformer with rotary position embedding. Neurocomputing, pp. 127063, 2023.
  41. 41.Theodorou, E., Buchli, J., and Schaal, S. A generalized path integral control approach to reinforcement learning. The Journal of Machine Learning Research, 11:3137–3181, 2010.
  42. 42.Theodorou, E. A. Nonlinear stochastic control and information theoretic dualities: Connections, interdependencies and thermodynamic interpretations. Entropy, 17(5):3352–3375, 2015.
  43. 43.Vargas, F., Grathwohl, W., and Doucet, A. Denoising diffusion samplers. arXiv preprint arXiv:2302.13834, 2023.
  44. 44.von Rutte, D., Biggio, L., Kilcher, Y., and Hofmann, T. ¨ Figaro: Controllable music generation using learned and expert features. In The Eleventh International Conference on Learning Representations, 2022.
  45. 45.Wang, Z., Chen, K., Jiang, J., Zhang, Y., Xu, M., Dai, S., Gu, X., and Xia, G. Pop909: A pop-song dataset for music arrangement generation. arXiv preprint arXiv:2008.07142, 2020.
  46. 46.Wu, S.-L. and Yang, Y.-H. Musemorphose: Full-song and fine-grained piano music style transfer with one transformer vae. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 31:1953–1967, 2023.
  47. 47.Yang, L.-C. and Lerch, A. On the evaluation of generative models in music. Neural Computing and Applications, 32(9):4773–4784, 2020.
  48. 48.Yang, L.-C., Chou, S.-Y., and Yang, Y.-H. Midinet: A convolutional generative adversarial network for symbolic-domain music generation. International Society for Music Information Retrieval, 2017.
  49. 49.Yin, Z., Reuben, F., Stepney, S., and Collins, T. Deep learning’s shallow gains: a comparative evaluation of algorithms for automatic music generation. Machine Learning, 112(5):1785–1822, 2023.
  50. 50.Zhang, C., Ren, Y., Zhang, K., and Yan, S. Sdmuse: Stochastic differential music editing and generation via hybrid representation. IEEE Transactions on Multimedia, 2023a.
  51. 51.Zhang, Q. and Chen, Y. Path integral sampler: a stochastic control approach for sampling. arXiv preprint arXiv:2111.15141, 2021.
  52. 52.Zhang, Q., Song, J., Huang, X., Chen, Y., and Liu, M.-Y. Diffcollage: Parallel generation of large content with diffusion models. CVPR, 2023b.

Citation

MLA
Huang, Y., et al. “Symbolic Music Generation with Non-Differentiable Rule Guided Diffusion”. arXiv, 2024, https://doi.org/10.48550/arxiv.2402.14285.
APA
Huang, Y., Ghatare, A., Liu, Y., Hu, Z., Zhang, Q., Sastry, C. S., Gururani, S., Oore, S., & Yue, Y. (2024). Symbolic Music Generation with Non-Differentiable Rule Guided Diffusion. arXiv. https://doi.org/10.48550/arxiv.2402.14285
Chicago
Huang, Y., A. Ghatare, Y. Liu, et al. 2024. “Symbolic Music Generation with Non-Differentiable Rule Guided Diffusion”. Preprint, ArXiv. https://doi.org/10.48550/arxiv.2402.14285.
Harvard
Huang, Y. et al. (2024) “Symbolic Music Generation with Non-Differentiable Rule Guided Diffusion”. arXiv. Available at: https://doi.org/10.48550/arxiv.2402.14285.
Vancouver
1. Huang Y, Ghatare A, Liu Y, Hu Z, Zhang Q, Sastry CS, Gururani S, Oore S, Yue Y (2024) Symbolic Music Generation with Non-Differentiable Rule Guided Diffusion. https://doi.org/10.48550/arxiv.2402.14285

BibTeX

@misc{https://doi.org/10.48550/arxiv.2402.14285,
  doi = {10.48550/ARXIV.2402.14285},
  url = {https://arxiv.org/abs/2402.14285},
  author = {Huang, Yujia and Ghatare, Adishree and Liu, Yuanzhe and Hu, Ziniu and Zhang, Qinsheng and Sastry, Chandramouli S and Gururani, Siddharth and Oore, Sageev and Yue, Yisong},
  keywords = {Sound (cs.SD), Machine Learning (cs.LG), Audio and Speech Processing (eess.AS), FOS: Computer and information sciences, FOS: Computer and information sciences, FOS: Electrical engineering, electronic engineering, information engineering, FOS: Electrical engineering, electronic engineering, information engineering},
  title = {Symbolic Music Generation with Non-Differentiable Rule Guided Diffusion},
  publisher = {arXiv},
  year = {2024},
  copyright = {arXiv.org perpetual, non-exclusive license}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/