Built independently by an author, for readers. Read the story and support ChapterPal

topic

transition probabilities (transition probability)

Transition probabilities represent the conditional likelihoods of a system switching from one state to another within a stochastic process. A transition probability measures the specific chance that a system currently in a given state will move to a designated destination state in the next step or over a defined time interval. In computational modeling and applied mathematics, these values are central to Markov processes, where the likelihood of entering a future state depends solely on the current state rather than the historical sequence of preceding events. Typically structured into a transition matrix, transition probabilities are essential for analyzing algorithmic state machines, simulating probabilistic environments, modeling reinforcement learning policies, and controlling Markov jump systems subject to discrete mode changes.

3 items

Estimating Instance-dependent Bayes-label Transition Matrix using a Deep Neural Network

Estimating Instance-dependent Bayes-label Transition Matrix using a Deep Neural Network

Shuo Yang, Erkun Yang, Bo Han, Yang Liu, Min Xu, Gang Niu, Tongliang Liu

Why you should read this

Proposes modeling instance-dependent label noise by estimating transitions from Bayes optimal labels to noisy labels using a deep neural network, shrinking the search space and improving classification accuracy on noisy datasets.

In label-noise learning, estimating the transition matrix is a hot topic as the matrix plays an important role in building statistically consistent classifiers. Traditionally, the transition from clean labels to noisy labels (i.e., clean-label transition matrix (CLTM)) has been widely exploited to learn a clean label classifier by employing the noisy data. Motivated by that classifiers mostly output Bayes optimal labels for prediction, in this paper, we study to directly model the transition from Bayes optimal labels to noisy labels (i.e., Bayes-label transition matrix (BLTM)) and learn a classifier to predict Bayes optimal labels. Note that given only noisy data, it is ill-posed to estimate either the CLTM or the BLTM. But favorably, Bayes optimal labels have less uncertainty compared with the clean labels, i.e., the class posteriors of Bayes optimal labels are one-hot vectors while those of clean labels are not. This enables two advantages to estimate the BLTM, i.e., (a) a set of examples with theoretically guaranteed Bayes optimal labels can be collected out of noisy data; (b) the feasible solution space is much smaller. By exploiting the advantages, we estimate the BLTM parametrically by employing a deep neural network, leading to better generalization and superior classification performance.

Added

2026-10-03

Nearly Minimax Optimal Reinforcement Learning for Linear Markov Decision Processes

Nearly Minimax Optimal Reinforcement Learning for Linear Markov Decision Processes

Jiafan He, Heyang Zhao, Dongruo Zhou, Quanquan Gu

OrganizationsUniversity of California, Los Angeles

Why you should read this

Presents LSVI-UCB++, the first computationally efficient reinforcement learning algorithm for linear Markov decision processes to achieve nearly minimax optimal regret by combining variance-aware weighted regression with a rare-switching policy update scheme.

We study reinforcement learning (RL) with linear function approximation. For episodic time-inhomogeneous linear Markov decision processes (linear MDPs) whose transition probability can be parameterized as a linear function of a given feature mapping, we propose the first computationally efficient algorithm that achieves the nearly minimax optimal regret Õ(d√H³K), where d is the dimension of the feature mapping, H is the planning horizon, and K is the number of episodes. Our algorithm is based on a weighted linear regression scheme with a carefully designed weight, which depends on a new variance estimator that (1) directly estimates the variance of the optimal value function, (2) monotonically decreases with respect to the number of episodes to ensure a better estimation accuracy, and (3) uses a rare-switching policy to update the value function estimator to control the complexity of the estimated value function class. Our work provides a complete answer to optimal RL with linear MDPs, and the developed algorithm and theoretical tools may be of independent interest.

Added

2026-10-02

Structured Denoising Diffusion Models in Discrete State-Spaces

Structured Denoising Diffusion Models in Discrete State-Spaces

Jacob Austin, Daniel D. Johnson, Jonathan Ho, Daniel Tarlow, Rianne van den Berg

OrganizationsGoogleMicrosoft

Why you should read this

Establishes the foundational mathematical framework for applying continuous diffusion logic to categorical language tokens through explicit discrete transition matrices.

Denoising diffusion probabilistic models (DDPMs) (Ho et al. 2020) have shown impressive results on image and waveform generation in continuous state spaces. Here, we introduce Discrete Denoising Diffusion Probabilistic Models (D3PMs), diffusion-like generative models for discrete data that generalize the multinomial diffusion model of Hoogeboom et al. 2021, by going beyond corruption processes with uniform transition probabilities. This includes corruption with transition matrices that mimic Gaussian kernels in continuous space, matrices based on nearest neighbors in embedding space, and matrices that introduce absorbing states. The third allows us to draw a connection between diffusion models and autoregressive and mask-based generative models. We show that the choice of transition matrix is an important design decision that leads to improved results in image and text domains. We also introduce a new loss function that combines the variational lower bound with an auxiliary cross entropy loss. For text, this model class achieves strong results on character-level text generation while scaling to large vocabularies on LM1B. On the image dataset CIFAR-10, our models approach the sample quality and exceed the log-likelihood of the continuous-space DDPM model.

Added

2026-02-25