keyword
Time-homogeneous linear mixture MDPs
A time-homogeneous linear mixture Markov decision process is a mathematical model for sequential decision-making in reinforcement learning where the probability of transitioning from a given state and action to the next state is represented as a linear combination of known feature vectors, and this transition mechanism remains constant across every step of an episode. In this formulation, the transition probability distribution is parameterized by the inner product of a known feature mapping of the current state, action, and candidate next state with an unknown parameter vector. The time-homogeneous property dictates that this parameter vector and the underlying system dynamics are stationary and do not vary with the time step or stage of the planning horizon, unlike time-inhomogeneous models where transition dynamics change at each step. This shared, stationary structure enables reinforcement learning algorithms to efficiently approximate value functions and generalize across large or continuous state-action spaces by learning a single set of transition parameters.
1 item

