topic
stochastic control
Stochastic control is a subfield of control theory and computer science focused on designing strategies to regulate and optimize dynamic systems that operate under uncertainty, noise, or random disturbances. Unlike deterministic approaches, it explicitly models unpredictability in system states, environmental conditions, and sensor observations using probabilistic frameworks such as Markov processes and stochastic differential equations. The primary objective is to compute control policies that optimize a specified performance criterion, such as minimizing expected operational costs or maximizing long-term rewards over time. Through mathematical tools like dynamic programming, the Hamilton-Jacobi-Bellman equation, and Markov decision processes, stochastic control provides the theoretical foundation for decision-making in robotics, reinforcement learning, autonomous navigation, telecommunications, and quantitative finance.
2 items

Symbolic Music Generation with Non-Differentiable Rule Guided Diffusion
Yujia Huang, Adishree Ghatare, Yuanzhe Liu, Ziniu Hu, Qinsheng Zhang, Chandramouli Shama Sastry, Siddharth Gururani, Sageev Oore, Yisong Yue
Why you should read this
Introduces Stochastic Control Guidance, a training-free method that enables pre-trained diffusion models to adhere to non-differentiable symbolic musical rules using only forward function evaluations.
We study the problem of symbolic music generation (e.g., generating piano rolls), with a technical focus on non-differentiable rule guidance. Musical rules are often expressed in symbolic form on note characteristics, such as note density or chord progression, many of which are non-differentiable which pose a challenge when using them for guided diffusion. We propose Stochastic Control Guidance (SCG), a novel guidance method that only requires forward evaluation of rule functions that can work with pre-trained diffusion models in a plug-and-play way, thus achieving training-free guidance for non-differentiable rules for the first time. Additionally, we introduce a latent diffusion architecture for symbolic music generation with high time resolution, which can be composed with SCG in a plug-and-play fashion. Compared to standard strong baselines in symbolic music generation, this framework demonstrates marked advancements in music quality and rule-based controllability, outperforming current state-of-the-art generators in a variety of settings. For detailed demonstrations, code and model checkpoints, please visit our project website.
Added
2026-10-04

Non-stationary Online Learning with Memory and Non-stochastic Control
Peng Zhao, Yu-Hu Yan, Yu-Xiang Wang, Zhi-Hua Zhou
Why you should read this
Develops a switching-cost-aware online ensemble method that achieves optimal dynamic policy regret for online convex optimization with memory and yields the first provably competitive gradient-based controller for non-stationary, non-stochastic control.
We study the problem of Online Convex Optimization (OCO) with memory, which allows loss functions to depend on past decisions and thus captures temporal effects of learning problems. In this paper, we introduce dynamic policy regret as the performance measure to design algorithms robust to non-stationary environments, which competes algorithms’ decisions with a sequence of changing comparators. We propose a novel algorithm for OCO with memory that provably enjoys an optimal dynamic policy regret in terms of time horizon, non-stationarity measure, and memory length. The key technical challenge is how to control the switching cost, the cumulative movements of player’s decisions, which is neatly addressed by a novel switching-cost-aware online ensemble approach equipped with a new meta-base decomposition of dynamic policy regret and a careful design of meta-learner and base-learner that explicitly regularizes the switching cost. The results are further applied to tackle non-stationarity in online non-stochastic control (Agarwal et al., 2019), i.e., controlling a linear dynamical system with adversarial disturbance and convex cost functions. We derive a novel gradient-based controller with dynamic policy regret guarantees, which is the first controller provably competitive to a sequence of changing policies for online non-stochastic control.
Added
2026-10-03
