Offline Reinforcement Learning with Value-based Episodic Memory
Xiaoteng MaYiqin YangHao HuJun YangChongjie ZhangQianchuan ZhaoBin LiangQihan Liu
Develops Value-based Episodic Memory, an offline reinforcement learning framework that combines expectile state-value learning with trajectory-based implicit planning to avoid out-of-distribution extrapolation errors and achieve superior performance on sparse-reward tasks.
Deploying machine learning to optimize decision-making in safety-critical and high-risk environments often requires learning exclusively from pre-collected, offline data rather than live trial and error. A primary barrier in offline learning is extrapolation error, where algorithms severely overestimate the potential rewards of actions not present in the historical dataset. Most current approaches attempt to counteract this instability using complex constraints, behavioral models, or artificial penalties. The article introduces and evaluates Value-based Episodic Memory, a novel framework designed to learn effective decision policies strictly within the bounds of historical data without requiring auxiliary behavioral or dynamic models.
The approach combines two core concepts: Expectile V-Learning and implicit memory-based planning. Instead of evaluating action-specific values that risk extrapolation errors on unseen actions, the framework learns direct state values. An expectile parameter balances conservative imitation of past behavior against the pursuit of optimal performance. The method then performs recursive planning along historical trajectories to enhance advantage estimations, training the final decision policy through standard regression techniques. The authors tested this method across continuous control benchmarks, including the D4RL suite spanning sparse-reward robotic navigation, complex robotic manipulation, and standard locomotion tasks.
The evaluation produced several key findings regarding algorithm performance and stability. First, the proposed framework achieved superior or competitive results compared to leading baseline algorithms across the majority of benchmark tasks. Second, performance gains were especially pronounced in complex, sparse-reward environments; in challenging maze navigation and robotic manipulation tasks, the method achieved success rates substantially higher than existing methods, many of which failed entirely. Third, the framework avoided the catastrophic value overestimation and training collapse that compromised competing action-value models on narrow human demonstration datasets. Finally, theoretical analysis confirmed that the approach is provably convergent and accelerates value learning without introducing systematic bias.
These findings indicate that decision-making models can be reliably trained on offline operational data with lower computational complexity and greater stability. By eliminating the need for complex behavioral generative models or penalty tuning, the framework reduces implementation risk and engineering overhead for data-driven optimization in physical systems. Organizations evaluating autonomous systems can achieve higher operational performance from imperfect or sparse demonstration logs without risking unstable behavior.
To apply these insights, technical teams should consider adopting state-value expectile architectures when training models from limited operational logs. Practitioners must balance the expectile parameter, setting conservative values for narrow or noisy datasets and higher values when data coverage is extensive. Further work should explore deploying this framework in real-world physical pilots beyond simulated benchmarks and evaluating its performance in highly stochastic environments, as the theoretical guarantees assume deterministic system dynamics.
- Paper: Offline Reinforcement Learning with Implicit Q-Learning, Ilya Kostrikov et al. (2021). Implicit Q-Learning introduces the use of statistical expectile regression to learn in-sample value functions without out-of-distribution action queries, serving as a direct foundation for Expectile V-Learning.
- Paper: Conservative Q-Learning for Offline Reinforcement Learning, Aviral Kumar et al. (2020). Conservative Q-Learning formalizes value conservatism to counter extrapolation error in offline RL, establishing the baseline paradigm that VEM improves upon by shifting focus from Q-functions to V-functions.
- Paper: D4RL: Datasets for Deep Data-Driven Reinforcement Learning, Justin Fu et al. (2020). This benchmark paper establishes the standardized D4RL suite and evaluation protocols for offline reinforcement learning that VEM uses to validate its empirical claims.
- Paper: Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems, Sergey Levine et al. (2020). This comprehensive tutorial outlines the foundational challenges of distributional shift, value overestimation, and conservatism in offline reinforcement learning addressed by VEM.
- Paper: Off-Policy Deep Reinforcement Learning without Exploration, Scott Fujimoto et al. (2018). This work identifies and formalizes extrapolation error in offline batch settings, providing essential motivation for in-support value learning frameworks.
- Paper: Stabilizing Off-Policy Q-Learning via Bootstrapping Error Reduction, Aviral Kumar et al. (2019). BEAR analyzes error propagation and enforces support constraints on policy updates in offline RL, motivating subsequent in-sample value estimation approaches.
- Paper: Self-improving reactive agents based on reinforcement learning, planning and teaching, Longxin Lin (1992). This seminal paper introduces experience replay and episodic memory mechanisms in reinforcement learning that inspire trajectory-based episodic value learning.
- Paper: Efficient Diffusion Policies For Offline Reinforcement Learning, Bingyi Kang et al. (2023). This work pairs diffusion models with in-sample implicit value estimators like those in offline expectile learning to enable fast, highly expressive policy extraction.
- Paper: Planning with Diffusion for Flexible Behavior Synthesis, Michael Janner et al. (2022). Diffuser advances offline planning by framing trajectory synthesis as conditional diffusion guidance, generalizing trajectory-based implicit planning concepts.
- Paper: Plan Better Amid Conservatism: Offline Multi-Agent Reinforcement Learning with Actor Rectification, Ling Pan et al. (2022). This paper extends conservative offline reinforcement learning paradigms into multi-agent decision-making environments via actor rectification.
- Paper: VIP: Towards Universal Visual Reward and Representation via Value-Implicit Pre-Training, Yecheng Jason Ma et al. (2023). VIP utilizes implicit value optimization principles from offline data to formulate self-supervised visual representation and reward learning.
- Paper: VIMPO: Value-Implicit Policy Optimization for LLMs, Zhewei Kang et al. (2026). VIMPO adapts implicit value formulations directly to verifiable reasoning tasks in large language models to avoid explicit learned critics.
- Paper: Reinforcement Learning: An Overview, Kevin P. Murphy (2024). This comprehensive survey provides an updated overview of reinforcement learning paradigms, contextualizing offline value methods within modern RL literature.
