Built independently by an author, for readers. Read the story and support ChapterPal

keyword

state-action value function

The state-action value function, commonly known as the Q-function in reinforcement learning, is a mathematical function that estimates the expected total return an agent will receive by taking a specific action in a given state and following a particular policy thereafter. It evaluates the long-term utility of an action by balancing the immediate reward against discounted future rewards accumulated across subsequent time steps. Formulated recursively through the Bellman equation, this function allows decision-making systems to compare different choices at any state and update their behavioral strategies to find optimal policies, serving as a core component of both value-based and actor-critic learning algorithms.

2 items

Model-Bellman Inconsistency for Model-based Offline Reinforcement Learning

Model-Bellman Inconsistency for Model-based Offline Reinforcement Learning

Yihao Sun, Jiaji Zhang, Chengxing Jia, Haoxin Lin, Junyin Ye, Yang Yu

Why you should read this

Proposes MOBILE, a model-based offline reinforcement learning algorithm that measures uncertainty using Bellman estimation inconsistency across a dynamics ensemble to approximate true Bellman errors and achieve state-of-the-art performance on D4RL and NeoRL benchmarks.

For offline reinforcement learning (RL), model-based methods are expected to be data-efficient as they incorporate dynamics models to generate more data. However, due to inevitable model errors, straightforwardly learning a policy in the model typically fails in the offline setting. Previous studies have incorporated conservatism to prevent out-of-distribution exploration. For example, MOPO penalizes rewards through uncertainty measures from predicting the next states, which we have discovered are loose bounds of the ideal uncertainty, i.e., the Bellman error. In this work, we propose MOdel-Bellman Inconsistency penalized OffLinE Policy Optimization (MOBILE), a novel uncertainty-driven offline RL algorithm. MOBILE conducts uncertainty quantification through the inconsistency of Bellman estimations under an ensemble of learned dynamics models, which can be a better approximator to the true Bellman error, and penalizes the Bellman estimation based on this uncertainty. Empirically we have verified that our proposed uncertainty quantification can be significantly closer to the true Bellman error than the compared methods. Consequently, MOBILE outperforms prior offline RL approaches on most tasks of D4RL and NeoRL benchmarks.

Added

2026-10-03

Offline Reinforcement Learning with Implicit Q-Learning

Offline Reinforcement Learning with Implicit Q-Learning

Ilya Kostrikov, Ashvin Nair, Sergey Levine

OrganizationsUniversity of California Berkeley

Why you should read this

Introduces Implicit Q-Learning (IQL), an offline reinforcement learning algorithm that avoids querying out-of-distribution actions by using expectile regression on state values, achieving state-of-the-art performance on D4RL benchmarks while enabling effective online fine-tuning.

Offline reinforcement learning requires reconciling two conflicting aims: learning a policy that improves over the behavior policy that collected the dataset, while at the same time minimizing the deviation from the behavior policy so as to avoid errors due to distributional shift. This trade-off is critical, because most current offline reinforcement learning methods need to query the value of unseen actions during training to improve the policy, and therefore need to either constrain these actions to be in-distribution, or else regularize their values. We propose an offline RL method that never needs to evaluate actions outside of the dataset, but still enables the learned policy to improve substantially over the best behavior in the data through generalization. The main insight in our work is that, instead of evaluating unseen actions from the latest policy, we can approximate the policy improvement step implicitly by treating the state value function as a random variable, with randomness determined by the action (while still integrating over the dynamics to avoid excessive optimism), and then taking a state conditional upper expectile of this random variable to estimate the value of the best actions in that state. This leverages the generalization capacity of the function approximator to estimate the value of the best available action at a given state without ever directly querying a Q-function with this unseen action. Our algorithm alternates between fitting this upper expectile value function and backing it up into a Q-function. Then, we extract the policy via advantage-weighted behavioral cloning. We dub our method implicit Q-learning (IQL). IQL demonstrates the state-of-the-art performance on D4RL, a standard benchmark for offline reinforcement learning. We also demonstrate that IQL achieves strong performance fine-tuning using online interaction after offline initialization.

Added

2026-09-24