Actor Prioritized Experience Replay
Baturay SaglamFurkan B. MutluDogan Can ÇiçekSuleyman S. Kozat
Explains why standard prioritized experience replay fails in off-policy actor-critic continuous control and introduces LA3P, a replay prioritization framework that improves policy gradient accuracy by training the actor network on low temporal-difference error transitions.
Deep reinforcement learning enables autonomous agents to learn complex decision-making through trial and error by storing and reusing past experiences from a replay buffer. While prioritizing experiences with high temporal-difference (TD) error—known as Prioritized Experience Replay (PER)—substantially boosts performance in discrete action tasks like video games, it consistently degrades performance in continuous control domains like robotics. In continuous settings, actor-critic architectures must be used, but standard prioritization frequently leads to unstable training and suboptimal policies.
The article investigates the root theoretical causes behind this failure and proposes a novel sampling framework, termed Loss-Adjusted Approximate Actor Prioritized Experience Replay (LA3P). The study evaluates how scheduling different experience distributions for the actor (policy) and critic (value estimator) networks, combined with corrected loss functions, can restore and enhance the benefits of prioritization in continuous control.
The authors develop a mathematical proof showing that high TD errors in the critic directly increase value estimation errors, causing the calculated actor policy gradients to diverge from the true optimal gradient. To solve this, the proposed LA3P framework introduces inverse prioritized sampling, which supplies the actor network with low TD error transitions that the critic understands reliably. To preserve theoretical consistency and learning stability between the networks, LA3P allocates a fraction of each batch (typically 50%) to uniformly sampled experiences shared by both actor and critic. Furthermore, the framework integrates modified loss functions—the Huber loss and Prioritized Approximate Loss—to eliminate outlier bias caused by standard mean-squared error updates.
Key findings show that LA3P significantly outperforms standard actor-critic baselines (Soft Actor-Critic and Twin Delayed DDPG), vanilla PER, and competing correction methods across challenging continuous control benchmarks (MuJoCo and Box2D). Pairwise statistical tests confirmed that these performance gains are statistically significant (p < 0.05) across most complex environments, such as HalfCheetah, Ant, Hopper, and Swimmer. Ablation studies demonstrated that the shared uniform sampling component is the single most critical factor for maintaining stability, and a 50-50 split between uniform and prioritized/inverse-prioritized sampling consistently yields peak performance without requiring environment-specific tuning.
These results provide a clear blueprint for engineering more stable and data-efficient continuous control systems. The findings demonstrate that actor-critic architectures cannot treat experience sampling identically for policy generation and value estimation. By mitigating policy gradient divergence and eliminating outlier bias, organizations deploying continuous reinforcement learning can achieve faster convergence and higher final performance, lowering the compute time and sample collection costs associated with training complex models.
Practitioners looking to implement prioritization in continuous control should adopt the LA3P sampling framework with a default uniform batch fraction of 0.5 and the specified Huber and approximate loss corrections. Future research should investigate alternative shared sampling distributions aimed at variance reduction and evaluate the approach across real-world robotic systems and broader continuous-action applications.
The findings are supported by comprehensive theoretical derivations and rigorous statistical benchmarking over multiple random seeds. A minor limitation is that LA3P incurs a higher computational runtime due to maintaining additional tree data structures for inverse sampling, though parallel single-instruction multiple-data (SIMD) operations on modern hardware substantially mitigate this overhead. Readers can have high confidence in the algorithm's performance advantages in simulated continuous control benchmarks.
- Paper: Prioritized Experience Replay, Tom Schaul et al. (2016). Introduces the foundational Prioritized Experience Replay (PER) mechanism using temporal-difference error that the target paper directly investigates and modifies for continuous actor-critic control.
- Paper: Addressing Function Approximation Error in Actor-Critic Methods, Scott Fujimoto et al. (2018). Presents Twin Delayed Deep Deterministic Policy Gradient (TD3), one of the primary continuous actor-critic baseline architectures evaluated and stabilized in the target work.
- Paper: Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor, Tuomas Haarnoja et al. (2018). Establishes Soft Actor-Critic (SAC), a core continuous actor-critic algorithm that serves as a primary benchmark and foundation for LA3P's sampling corrections.
- Paper: Continuous control with deep reinforcement learning, T. Lillicrap et al. (2015). Provides the foundational deep deterministic policy gradient framework for continuous control that underpins modern deep actor-critic methods and their experience replay dynamics.
- Paper: Deterministic Policy Gradient Algorithms, David Silver et al. (2014). Derives the theoretical foundations of deterministic policy gradients and off-policy actor-critic architectures analyzed in the target paper's policy gradient divergence proofs.
- Paper: Actor-Critic Algorithms, Vijay Konda et al. (1999). Establishes the two-time-scale convergence framework between actor policy updates and critic value estimation that informs how distinct sampling distributions affect learning stability.
- Paper: Self-improving reactive agents based on reinforcement learning, planning and teaching, Longxin Lin (1992). Introduces the fundamental concept and utility of experience replay in connectionist reinforcement learning.
No sufficiently relevant recommendations were found.
