Retrieval-Augmented Reinforcement Learning
Anirudh GoyalAbram L. FriesenAndrea BaninoTheophane WeberNan Rosemary KeAdrià Puigdomènech BadiaArthur GuezMehdi MirzaPeter Conway HumphreysKsenia Konyushkova
Proposes an architecture that augments reinforcement learning agents with an attention-based neural retrieval process, allowing policies to dynamically query past trajectories to reduce task interference in multi-task offline settings and significantly accelerate learning in online Atari benchmarks.
Modern artificial intelligence systems using reinforcement learning typically compress all previous experience into fixed network parameters via repetitive gradient updates. This standard paradigm faces critical limitations: it demands excessive computation, requires numerous updates to integrate new experiences, and restricts agent behavior based on fixed model capacity. Consequently, agents frequently struggle to recall specific, relevant past situations when facing complex or multi-task scenarios.
The article demonstrates an alternative approach called Retrieval-Augmented Reinforcement Learning, which pairs an agent with a dedicated retrieval mechanism that dynamically fetches relevant contextual information from external datasets of past trajectories during decision-making.
The authors evaluate this framework across standard video game benchmarks (Atari) and complex multi-task offline environments (Gridroboman, BabyAI, and CausalWorld continuous control). The retrieval process uses attention mechanisms over summarized past experiences and regulates information exchange using an information bottleneck to avoid overload. Crucially, the retrieval process and the decision-making agent are parameterized as separate interacting components.
The findings establish that retrieval augmentation significantly enhances agent efficiency and capabilities. On Atari games, the retrieval-augmented agent achieved an 11.32% improvement in mean human-normalized score compared to a leading baseline, demonstrating particular strength in tasks requiring extended planning. In offline multi-task environments, retrieval-augmented models prevented task interference as the number of tasks scaled from 10 to 30, where standard architectures degraded. On compositional language-directed benchmarks, retrieval increased task success rates from 45% to 74% with multi-task data, as the architecture successfully reused sub-task experiences.
These results demonstrate that externalizing memory retrieval relieves network capacity constraints, reduces data collection burdens, and improves performance in multi-task operations. This shift suggests organizations can deploy more agile and data-efficient decision models by decoupling memory access from core parametric networks.
Organizations evaluating reinforcement learning for complex multi-task or robotic control should consider retrieval-augmented architectures over scaling model parameters alone. Future efforts should test retrieval across decentralized multi-agent deployments and validate adaptations in zero-shot or few-shot environments.
Confidence in these findings is high for simulated control and standard benchmarks across multiple tested random seeds. However, decision-makers should exercise caution because computational overhead increases when querying large trajectory batches, and the methodology has not yet been demonstrated in real-world physical systems or massive multi-agent competitive environments.
- Paper: Prioritized Experience Replay, Tom Schaul et al. (2016). Prioritized Experience Replay shows how selecting informative past transitions can improve RL learning, a useful precursor to retrieving relevant experience for decisions.
- Paper: Experience Replay for Continual Learning, David Rolnick et al. (2018). Experience Replay for Continual Learning establishes how replaying stored experience can protect performance across sequential tasks, grounding the source’s concern with multi-task interference.
- Paper: Deep Recurrent Q-Learning for Partially Observable MDPs, Matthew Hausknecht et al. (2015). Deep Recurrent Q-Learning explains how RL agents use memory to exploit past observations, clarifying the alternative to the source’s externalized retrieval.
- Paper: Meta-Learning with Memory-Augmented Neural Networks, Adam Santoro et al. (2016). Meta-Learning with Memory-Augmented Neural Networks develops content-based access to external memory, a key conceptual foundation for retrieving information separately from model parameters.
- Paper: Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning, Bowen Jin et al. (2025). Search-R1 extends retrieval-augmented decision-making into language-model RL, training agents to make and use multi-step external searches through reinforcement learning.
