Symbolic Music Generation with Non-Differentiable Rule Guided Diffusion
Yujia HuangAdishree GhatareYuanzhe LiuZiniu HuQinsheng ZhangChandramouli Shama SastrySiddharth GururaniSageev OoreYisong Yue
Introduces Stochastic Control Guidance, a training-free method that enables pre-trained diffusion models to adhere to non-differentiable symbolic musical rules using only forward function evaluations.
Generating symbolic music (such as piano rolls) with modern generative models offers significant potential for human-AI co-creation. However, giving human composers precise control over specific musical rules—such as target chord progressions or note densities—remains a fundamental challenge. Existing diffusion-based models require differentiable loss functions to calculate gradients for guidance or necessitate training task-specific surrogate models, which are resource-intensive, difficult to train accurately for complex rules, and impractical for flexible, multi-rule compositions.
The article introduces and evaluates Stochastic Control Guidance (SCG), a training-free, plug-and-play guidance framework that controls pre-trained diffusion models using non-differentiable and black-box rules. It pairs this method with a high-resolution latent diffusion architecture designed to capture expressive piano performances with 10-millisecond timing precision.
To achieve non-differentiable control, the authors formulate the generation process as a stochastic optimal control problem. Rather than calculating mathematical gradients, SCG generates multiple candidate realizations at each intermediate generation step, estimates the final clean output for each candidate, and selects the direction that minimizes rule violation based purely on forward evaluation. The system was trained on approximately 1,000 hours of diverse classical and pop piano datasets and systematically evaluated across unconditional generation, individual rule guidance, composite multi-rule guidance, and music editing tasks against competitive baselines.
The findings show that SCG provides superior controllability over non-differentiable musical rules. In individual rule testing on non-differentiable note density, SCG lowered the error metric from 2.486 (unconditional) down to 0.131, markedly outperforming surrogate classifier methods (0.698) and neural-network-based loss guidance (1.261). For chord progression adherence, SCG reduced the error metric to 0.273, compared to 0.723 for classifier guidance. In composite multi-rule settings, combining SCG with lightweight classifier guidance produced the best overall rule adherence across all target rules simultaneously while maintaining high musical fidelity. Additionally, human listening surveys with 15 experienced musicians ranked the proposed system highest across rule alignment, musical creativity, coherence, and overall aesthetic appeal.
These results demonstrate that generative models can achieve fine-grained symbolic control without retraining or relying on fragile differentiable surrogates. In practice, this eliminates the high compute costs associated with maintaining custom models for every desired musical rule. It also enables broader applications across fields where non-differentiable constraints or black-box physics simulators must guide generative processes.
For practical implementation, practitioners and creators can adopt SCG as a plug-and-play compositional tool. When low latency is required, teams should combine coarse classifier guidance with a modest candidate sample size (e.g., four samples per step) or apply stochastic fast-sampling algorithms, which achieve comparable control while speeding up generation roughly fourfold.
The primary limitation of SCG is its increased computational cost during sampling due to generating and evaluating multiple candidates per step. Users must also manage an inherent trade-off where excessively optimizing a single constraint can degrade overall musical coherence. Nevertheless, the underlying theoretical formulation and extensive empirical validations provide strong confidence in SCG as an effective method for non-differentiable controlled generation.
- Paper: Loss-Guided Diffusion Models for Plug-and-Play Controllable Generation, Jiaming Song et al. (2023). Its plug-and-play loss guidance provides the differentiable-control approach that SCG replaces with forward evaluation for non-differentiable rules.
- Paper: Score-Based Generative Modeling through Stochastic Differential Equations, Yang Song et al. (2021). Its SDE formulation of score-based generation supplies the diffusion-process foundation needed to follow the source’s sampling-time guidance method.
- Paper: Denoising Diffusion Probabilistic Models, Jonathan Ho et al. (2020). Its DDPM account establishes the iterative denoising process that SCG steers during generation.
- Paper: Classifier-Free Diffusion Guidance, Jonathan Ho et al. (2022). Its classifier-free guidance method clarifies a standard diffusion-steering baseline that the source compares and combines with SCG.
No sufficiently relevant recommendations were found.
