Tell me why! Explanations support learning relational and causal structure
Andrew K. LampinenNicholas A. RoyIshita DasguptaStephanie C. Y. ChanAllison C. TamJames L. McClellandChen YanAdam SantoroNeil C. RabinowitzJane X. Wang
Demonstrates that training deep reinforcement learning agents to predict auxiliary language explanations enables them to master complex relational reasoning, resolve causal confounds, and perform experimental interventions to generalize to novel settings.
Artificial intelligence systems trained through reinforcement learning frequently struggle to acquire abstract relational concepts and causal structures from raw sensory inputs. While humans naturally rely on language explanations to highlight abstract principles and resolve ambiguous learning scenarios, standard machine learning agents tend to latch onto superficial shortcut features, leading to poor out-of-distribution generalization and failure in complex environments.
The article evaluates whether training reinforcement learning agents to generate natural language descriptions and explanations as an internal learning target enables them to master relational reasoning, disentangle confounded causal variables, and perform active experimental interventions.
To evaluate this capability, the researchers conducted simulated experiments across two-dimensional and three-dimensional environments using visual "odd-one-out" tasks. In these environments, neural network agents observed visual inputs and received rewards for identifying unique objects across multiple varying feature dimensions, such as color, shape, size, and texture. The core methodological approach trained agents to predict context- and behavior-relevant language explanations—specifically property descriptions of encountered objects and explanatory feedback on reward outcomes—as an auxiliary learning objective during training, without requiring any language inputs or explanations at test time.
The findings show that auxiliary explanation prediction dramatically improves agent performance and generalization across all experimental conditions. In baseline odd-one-out tasks, agents trained with explanations achieved over 90% accuracy (91.3% in two-dimensional and 92.7% in three-dimensional settings), whereas agents trained without explanations achieved 61.9% in two dimensions and collapsed to near-chance levels (29.5%) in three dimensions. In ambiguous settings where multiple features were perfectly correlated during training, explanations targeting a single dimension guided agents to generalize along that specific dimension over 85% of the time during deconfounded evaluations, overcoming the baseline bias toward easy superficial features like color. Furthermore, in meta-learning tasks requiring active causal intervention, only agents predicting explanations learned to perform experiments to deduce the underlying causal rules, achieving 96.9% accuracy on easier levels and 90.5% on harder levels compared to approximately 24.6% without explanations. Control analyses confirmed that explanations must be dynamically tied to the agent's behavior and situational context to be effective, and that language prediction was learned rapidly before overall task mastery.
These results demonstrate that language prediction functions as a highly effective learning scaffold, shaping internal representations toward reusable causal and relational abstractions. By mitigating the risk of models relying on fragile shortcut features, this technique provides a practical mechanism to enhance decision-making reliability, sample efficiency, and safety in autonomous systems operating in partially observable or confounded environments.
Organizations developing autonomous agents for complex or mission-critical tasks should consider integrating auxiliary explanation-generation targets into their model training pipelines. Where programmatic ground truth for explanations is unavailable, teams can evaluate leveraging human annotations or pre-trained language captioning models to supply supervisory descriptions. Future initiatives should pilot these techniques across broader and less structured domains.
The primary limitation of the work is that the experiments relied on synthetic explanations within structured simulated environments rather than real-world tasks or open-domain human language. Consequently, while confidence is high that auxiliary explanation prediction strongly enhances relational and causal learning in structured settings, decision-makers should exercise caution and conduct domain-specific pilot studies before applying these methods to unstructured, open-ended environments where ground truth causal explanations may be unavailable.
- Paper: Reinforcement Learning with Unsupervised Auxiliary Tasks, Max Jaderberg et al. (2017). Read this foundational demonstration that auxiliary prediction objectives can improve reinforcement-learning representations before considering explanations as a language-based learning scaffold.
- Paper: Guiding Pretraining in Reinforcement Learning with Large Language Models, Yuqing Du et al. (2023). This later work carries language-guided learning into open-ended RL pretraining, extending the source’s case for language as a useful signal beyond the main reward.
- Paper: Learning to Model the World With Language, Jessy Lin et al. (2024). Building on language as a predictive learning target, this later work uses multimodal world modeling to help agents ground varied language in predicted observations and rewards.
