Tell me why! Explanations support learning relational and causal structure

Andrew K. LampinenNicholas A. RoyIshita DasguptaStephanie C. Y. ChanAllison C. TamJames L. McClellandChen YanAdam SantoroNeil C. RabinowitzJane X. Wang

article2022ICML54 citations

Demonstrates that training deep reinforcement learning agents to predict auxiliary language explanations enables them to master complex relational reasoning, resolve causal confounds, and perform experimental interventions to generalize to novel settings.

Listen

Artificial intelligence systems trained through reinforcement learning frequently struggle to acquire abstract relational concepts and causal structures from raw sensory inputs. While humans naturally rely on language explanations to highlight abstract principles and resolve ambiguous learning scenarios, standard machine learning agents tend to latch onto superficial shortcut features, leading to poor out-of-distribution generalization and failure in complex environments.

The article evaluates whether training reinforcement learning agents to generate natural language descriptions and explanations as an internal learning target enables them to master relational reasoning, disentangle confounded causal variables, and perform active experimental interventions.

To evaluate this capability, the researchers conducted simulated experiments across two-dimensional and three-dimensional environments using visual "odd-one-out" tasks. In these environments, neural network agents observed visual inputs and received rewards for identifying unique objects across multiple varying feature dimensions, such as color, shape, size, and texture. The core methodological approach trained agents to predict context- and behavior-relevant language explanations—specifically property descriptions of encountered objects and explanatory feedback on reward outcomes—as an auxiliary learning objective during training, without requiring any language inputs or explanations at test time.

The findings show that auxiliary explanation prediction dramatically improves agent performance and generalization across all experimental conditions. In baseline odd-one-out tasks, agents trained with explanations achieved over 90% accuracy (91.3% in two-dimensional and 92.7% in three-dimensional settings), whereas agents trained without explanations achieved 61.9% in two dimensions and collapsed to near-chance levels (29.5%) in three dimensions. In ambiguous settings where multiple features were perfectly correlated during training, explanations targeting a single dimension guided agents to generalize along that specific dimension over 85% of the time during deconfounded evaluations, overcoming the baseline bias toward easy superficial features like color. Furthermore, in meta-learning tasks requiring active causal intervention, only agents predicting explanations learned to perform experiments to deduce the underlying causal rules, achieving 96.9% accuracy on easier levels and 90.5% on harder levels compared to approximately 24.6% without explanations. Control analyses confirmed that explanations must be dynamically tied to the agent's behavior and situational context to be effective, and that language prediction was learned rapidly before overall task mastery.

These results demonstrate that language prediction functions as a highly effective learning scaffold, shaping internal representations toward reusable causal and relational abstractions. By mitigating the risk of models relying on fragile shortcut features, this technique provides a practical mechanism to enhance decision-making reliability, sample efficiency, and safety in autonomous systems operating in partially observable or confounded environments.

Organizations developing autonomous agents for complex or mission-critical tasks should consider integrating auxiliary explanation-generation targets into their model training pipelines. Where programmatic ground truth for explanations is unavailable, teams can evaluate leveraging human annotations or pre-trained language captioning models to supply supervisory descriptions. Future initiatives should pilot these techniques across broader and less structured domains.

The primary limitation of the work is that the experiments relied on synthetic explanations within structured simulated environments rather than real-world tasks or open-domain human language. Consequently, while confidence is high that auxiliary explanation prediction strongly enhances relational and causal learning in structured settings, decision-makers should exercise caution and conduct domain-specific pilot studies before applying these methods to unstructured, open-ended environments where ground truth causal explanations may be unavailable.

  • Paper: Reinforcement Learning with Unsupervised Auxiliary Tasks, Max Jaderberg et al. (2017). Read this foundational demonstration that auxiliary prediction objectives can improve reinforcement-learning representations before considering explanations as a language-based learning scaffold.
Cover for Tell me why! Explanations support learning relational and causal structure

Abstract

Inferring the abstract relational and causal structure of the world is a major challenge for reinforcement-learning (RL) agents. For humans, language—particularly in the form of explanations—plays a considerable role in overcoming this challenge. Here, we show that language can play a similar role for deep RL agents in complex environments. While agents typically struggle to acquire relational and causal knowledge, augmenting their experience by training them to predict language descriptions and explanations can overcome these limitations. We show that language can help agents learn challenging relational tasks, and examine which aspects of language contribute to its benefits. We then show that explanations can help agents to infer not only relational but also causal structure. Language can shape the way that agents to generalize out-of-distribution from ambiguous, causally-confounded training, and explanations even allow agents to learn to perform experimental interventions to identify causal relationships. Our results suggest that language description and explanation may be powerful tools for improving agent learning and generalization.

Table of Contents

  • 1. The odd-one-out tasks
  • 2. Method: generating explanations
  • 3. Experiments
  • 3.1. Odd-one-out tasks in 2D and 3D RL environments
  • 3.2. Explanations can deconfound
  • 3.3. Explanations allow agents to learn to experiment
  • 3.4. Exploring the benefits of explanation in more detail
  • 4. Related work
  • 4.1. Related work in AI
  • 5. Discussion
  • Acknowledgments
  • References
  • A. Ablation experiments & further analyses
  • A.1. Agents trained without explanations fixate on the easiest feature dimensions
  • A.2. Explanations are most useful if they engage with the agent's behavior; shuffled explanations are useless
  • A.3. Providing explanations as agent inputs is not beneficial, and interferes with learning from explanation targets
  • A.4. Different kinds of explanations have complementary, sometimes separable benefits
  • A.5. The benefits of explanations depend on task complexity
  • A.6. Language prediction is learned faster than RL
  • A.7. Learning properties through a curriculum rather than auxiliary losses
  • A.8. Auxiliary unsupervised losses are neither necessary nor sufficient; thus the benefits of explanations are not simply due to more supervision
  • B. Quantitative results
  • C. Methods
  • C.1. RL agents & training
  • C.2. RL environment details
  • C.2.1. 2D

Knowls

  1. Knowl 1 — Task-relevant language explanations connect situations, actions, and abstract structure

    definition

    A language explanation is an utterance that connects a particular situation and the agent’s behavior to abstract, task-relevant structure. In the odd-one-out tasks, property explanations are given while the agent encounters an object and describe its features; reward explanations follow the agent’s choice and identify the feature or features that made the choice correct or incorrect. A statement about an object that does not relate to the task or the agent’s behavior—for example, an irrelevant visual attribute—does not meet this definition.

  2. Knowl 2 — Explanation prediction is an auxiliary target for an IMPALA agent

    model/method

    The agent receives visual observations and acts using an IMPALA reinforcement-learning setup with V-trace. A visual CNN in the 2D tasks or ResNet in the 3D tasks encodes each observation; a four-layer Gated Transformer-XL memory processes the visual representation and previous reward. Policy and value heads are multilayer perceptrons, and a one-layer LSTM explanation head predicts a language sequence from the memory state. Explanation predictions are trained with summed word-level softmax cross-entropy. Explanations are generated online conditional on the situation and agent behavior, and serve as auxiliary prediction targets rather than direct behavioral instructions or agent inputs. The policy is evaluated without explanations. The main 2D experiments also used an image-reconstruction auxiliary loss; reconstruction was disabled in the 3D experiments.

  3. Knowl 3 — Explanations improve relational odd-one-out learning in 2D and 3D

    empirical result

    In the odd-one-out task, an agent must select one of four objects that is unique along one feature dimension, while attributes on the other dimensions occur in pairs. The 2D environment uses color, texture, shape, and position; the 3D environment uses color, texture, size, and position, and its restricted first-person view can require remembering objects to compare them. Across evaluations in the final 1% of training, agents trained to predict property and reward explanations achieved 91.3 ± 0.7% accuracy in 2D and 92.7 ± 1.4% in 3D. Agents without explanations achieved 61.9 ± 2.2% in 2D and 29.5 ± 0.7% in 3D. The 2D results averaged five seeds per condition; the 3D results averaged three. In the 2D task, agents without explanations learned the easiest dimensions (position, and partly color) more readily than shape and texture; explanation-trained agents learned across dimensions. The authors’ analysis attributes the no-explanation gap to a tendency to rely on easy features rather than mastering all relevant relations.

  4. Knowl 4 — Single-dimension explanations shape out-of-distribution choices under confounding

    empirical result

    During confounded training, the target object was simultaneously unique in color, shape, and texture, so any of those perfectly correlated dimensions could predict reward. During deconfounded evaluation, different objects were unique in color, shape, and texture, requiring the agent’s choice to reveal which dimension it had learned to use. Without explanations, agents selected the color-unique object 55.4 ± 2.6% of the time, the shape-unique object 24.2 ± 7.6%, and the texture-unique object 15.4 ± 6.7%. Training with explanations consistently naming just one dimension shifted choices toward that dimension: choices matched color explanations 95.5 ± 0.9% of the time, shape explanations 87.5 ± 2.9%, and texture explanations 86.2 ± 0.9%. Thus, explanations altered which feature the agent used to generalize, despite not changing the features’ correlations with reward during training. Results averaged three seeds per condition.

  5. Knowl 5 — Agents learn causal experiments in a four-trial intervention task

    empirical result

    Each meta-learning episode contains four odd-one-out trials with a single reward-relevant dimension, such as color, that varies between episodes and is not directly observed. On the first three trials, the agent can use a wand once per trial to change the color, shape, or texture of an adjacent object; it must use the resulting outcome and reward feedback to infer the relevant dimension. Initial object attributes are either all the same (the easier condition) or paired (the harder condition), where changing one object can make a different object unique. A final deconfounded trial has a distinct unique object for each dimension, disables the wand, and awards 10 for a correct choice; correct choices on each earlier trial award 1. Explanations identify the episode’s relevant dimension. With explanations, agents reached 96.9 ± 0.3% final-choice accuracy on the easy condition and 90.5 ± 1.6% on the hard condition; without explanations they reached 24.0 ± 0.6% and 24.6 ± 1.2%, respectively. Results averaged four seeds per condition. The findings show that explanation-trained agents learned to use interventions to identify task-specific causal structure in this setting.

  6. Knowl 6 — Behavior- and context-relevant explanations are most useful

    empirical result

    The experiments distinguish explanations that respond to the agent’s behavior and current situation from signals that merely describe a possible scene or are unrelated to it. In a behavior-irrelevant condition, an explanation possible in the current room was supplied on approximately 10% of steps regardless of the agent’s actions. Such explanations supported some learning on the basic 2D task but were less effective and slower than behavior-relevant explanations. In the more demanding causal-intervention task, learning required explanations that were relevant to both the situation and the agent’s behavior. Explanations sampled without regard to the current context did not benefit learning. These comparisons indicate that task-relevant information alone is not always enough: in challenging settings, the explanation’s relationship to the agent’s actual behavior matters.

  7. Knowl 7 — Explanation targets outperform explanation inputs and generic auxiliary signals

    empirical result

    Training the agent to predict explanations was more effective than supplying explanations as inputs. Input explanations alone were not beneficial compared with explanation prediction, and combining explanation inputs with explanation targets impaired learning relative to using targets alone. In the 2D odd-one-out task, image reconstruction was neither necessary for explanation benefits nor sufficient to produce them without explanations. A separate curriculum that trained the agent to identify instructed object properties also failed to improve odd-one-out learning over the no-explanation baseline, whether it used a shared policy or separate policies. Together, these controls support the interpretation that the benefit comes from predicting task-relevant explanations, not simply from additional input information, more supervision, or learning object properties in isolation.

  8. Knowl 8 — Property and reward explanations have complementary benefits

    empirical result

    The relative usefulness of property and reward explanations depended on the task. In the basic 2D odd-one-out task, either type alone supported learning, but providing both led to faster learning. In the visually and memory-challenging 3D task, property explanations were particularly beneficial, while reward explanations alone were less effective. In the causal-intervention task, reward explanations were needed for learning when provided alone; property explanations alone did not support learning, while combining both types produced more complete learning within the training budget. The results suggest that descriptions of encountered objects and feedback about the consequences of choices supply different, complementary signals.

  9. Knowl 9 — Explanation prediction is learned before task mastery

    empirical result

    In the 2D odd-one-out experiments, language-prediction loss decreased substantially early in training, before the agent had mastered the reinforcement-learning task. Property explanations were learned especially quickly; reward explanations were learned more slowly, consistent with their dependence on understanding more of the task structure. The authors interpret this timing as consistent with explanations making task-relevant abstractions easier to learn and thereby supporting later task learning; the timing itself does not establish that causal mechanism.

  10. Knowl 10 — The evidence is limited to synthetic explanations and odd-one-out task variants

    limitation

    The experiments use variations of odd-one-out tasks in 2D and 3D environments, with explanations synthetically generated from known task structure. The results therefore do not establish that the same benefits hold in other domains or when ground-truth explanations are unavailable. Agents also learned the language patterns alongside the tasks through repeated examples; unlike humans, they were not shown to learn from a single explanation using prior language knowledge. The authors note that the tasks could in principle be learned from reward with sufficient data or other methods, and that enforcing simple explanations could be detrimental in domains that are difficult to explain or irreducibly complex.

Coverage note — Omitted detailed environment and optimizer hyperparameters because they do not add comparable significance to the main contributions; the central task designs, explanation-training approach, findings, controls, and stated limitations are included.

References

  1. 1.Ahn, W.-k., Brewer, W. F., and Mooney, R. J. Schema acquisition from a single example. Journal of Experimental Psychology: Learning, Memory, and Cognition, 18(2):391, 1992.
  2. 2.Andreas, J., Klein, D., and Levine, S. Learning with latent language. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pp. 2166–2179, 2018.
  3. 3.Bradbury, J., Frostig, R., Hawkins, P., Johnson, M. J., Leary, C., Maclaurin, D., Necula, G., Paszke, A., VanderPlas, J., Wanderman-Milne, S., and Zhang, Q. JAX: composable transformations of Python+NumPy programs, 2018. URL http://github.com/google/jax.
  4. 4.Brenner, A., Maurin, A.-S., Skiles, A., Stenwall, R., and Thompson, N. Metaphysical Explanation. In Zalta, E. N. (ed.), The Stanford Encyclopedia of Philosophy. Metaphysics Research Lab, Stanford University, Winter 2021 edition, 2021.
  5. 5.Cabi, S., Colmenarejo, S. G., Novikov, A., Konyushkova, K., Reed, S., Jeong, R., Zolna, K., Aytar, Y., Budden, D., Vecerik, M., et al. Scaling data-driven robotics with reward sketching and batch reinforcement learning. arXiv preprint arXiv:1909.12200, 2019.
  6. 6.Camburu, O.-M., Rocktaschel, T., Lukasiewicz, T., and Blunsom, P. e-snli: Natural language inference with natural language explanations. Advances in Neural Information Processing Systems, 31:9539–9549, 2018.
  7. 7.Cassens, J., Habenicht, L., Blohm, J., Wegener, R., Korman, J., Khemlani, S., Gronchi, G., Byrne, R. M., Warren, G., Quinn, M. S., et al. Explanation in human thinking. In Proceedings of the Annual Meeting of the Cognitive Science Society, volume 43, 2021.
  8. 8.Chen, J., Song, L., Wainwright, M., and Jordan, M. Learning to explain: An information-theoretic perspective on model interpretation. In International Conference on Machine Learning, pp. 883–892. PMLR, 2018.
  9. 9.Chi, M. T., De Leeuw, N., Chiu, M.-H., and Lavancher, C. Eliciting self-explanations improves understanding. Cognitive Science, 18(3):439–477, 1994. ISSN 0364-0213. doi: https://doi.org/10.1016/0364-0213(94)90016-7. URL https://www.sciencedirect.com/science/article/pii/0364021394900167.
  10. 10.Crutch, S. J., Connell, S., and Warrington, E. K. The different representational frameworks underpinning abstract and concrete knowledge: Evidence from odd-one-out judgements. Quarterly Journal of Experimental Psychology, 62(7):1377–1390, 2009.
  11. 11.Dasgupta, I. and Gershman, S. J. Memory as a computational resource. Trends in Cognitive Sciences, 2021.
  12. 12.Dasgupta, I., Wang, J., Chiappa, S., Mitrovic, J., Ortega, P., Raposo, D., Hughes, E., Battaglia, P., Botvinick, M., and Kurth-Nelson, Z. Causal reasoning from meta-reinforcement learning. arXiv preprint arXiv:1901.08162, 2019.
  13. 13.Dove, G. More than a scaffold: Language is a neuroenhancement. Cognitive Neuropsychology, 37 (5-6):288–311, 2020. doi: 10.1080/02643294.2019.1637338. URL https://doi.org/10.1080/02643294.2019.1637338. PMID: 31269862.
  14. 14.Edmiston, P. and Lupyan, G. What makes words special? words as unmotivated cues. Cognition, 143:93–100, 2015.
  15. 15.Edwards, B. J., Williams, J. J., Gentner, D., and Lombrozo, T. Explanation recruits comparison in a category-learning task. Cognition, 185:21–38, 2019.
  16. 16.Espeholt, L., Soyer, H., Munos, R., Simonyan, K., Mnih, V., Ward, T., Doron, Y., Firoiu, V., Harley, T., Dunning, I., et al. Impala: Scalable distributed deep-rl with importance weighted actor-learner architectures. In International Conference on Machine Learning, pp. 1407–1416. PMLR, 2018.
  17. 17.Fodor, J. A. and Pylyshyn, Z. W. Connectionism and cognitive architecture: A critical analysis. Cognition, 28(1-2):3–71, 1988.
  18. 18.Fyfe, E. R., McNeil, N. M., Son, J. Y., and Goldstone, R. L. Concreteness fading in mathematics and science instruction: A systematic review. Educational psychology review, 26(1):9–25, 2014.
  19. 19.Garcez, A. d. and Lamb, L. C. Neurosymbolic ai: the 3rd wave. arXiv preprint arXiv:2012.05876, 2020.
  20. 20.Geiger, A., Carstensen, A., Frank, M. C., and Potts, C. Relational reasoning and generalization using non-symbolic neural networks. arXiv preprint arXiv:2006.07968, 2020.
  21. 21.Geirhos, R., Jacobsen, J.-H., Michaelis, C., Zemel, R., Brendel, W., Bethge, M., and Wichmann, F. A. Shortcut learning in deep neural networks. Nature Machine Intelligence, 2(11):665–673, 2020.
  22. 22.Gentner, D. Why we’re so smart. Language in mind: Advances in the study of language and thought, 195235, 2003.
  23. 23.Gentner, D. and Christie, S. Relational language supports relational cognition in humans and apes. Behavioral and Brain Sciences, 31(2):136–137, 2008.
  24. 24.Ghosh, D., Rahme, J., Kumar, A., Zhang, A., Adams, R. P., and Levine, S. Why generalization in rl is difficult: Epistemic pomdps and implicit partial observability. Advances in Neural Information Processing Systems, 34, 2021.
  25. 25.Gopnik, A. and Sobel, D. M. Detecting blickets: How young children use information about novel causal powers in categorization and induction. Child development, 71(5):1205–1222, 2000.
  26. 26.Gopnik, A., Meltzoff, A. N., and Kuhl, P. K. The scientist in the crib: Minds, brains, and how children learn. William Morrow & Co, 1999.
  27. 27.Goyal, P., Niekum, S., and Mooney, R. J. Using natural language for reward shaping in reinforcement learning. arXiv preprint arXiv:1903.02020, 2019.
  28. 28.Gregor, K., Jimenez Rezende, D., Besse, F., Wu, Y., Merzic, H., and van den Oord, A. Shaping belief states with generative environment models for rl. In Advances in Neural Information Processing Systems, volume 32. Curran Associates, Inc., 2019. URL https://proceedings.neurips.cc/paper/2019/file/2c048d74b3410237704eb7f93a10c9d7-Paper.pdf.
  29. 29.Guan, L., Verma, M., Guo, S., Zhang, R., and Kambhampati, S. Widening the pipeline in human-guided reinforcement learning with explanation and context-aware data augmentation. In Advances in Neural Information Processing Systems, 2021.
  30. 30.Hase, P. and Bansal, M. When can models learn from explanations? a formal framework for understanding the roles of explanation data. arXiv preprint arXiv:2102.02201, 2021.
  31. 31.Hennigan, T., Cai, T., Norman, T., and Babuschkin, I. Haiku: Sonnet for JAX, 2020. URL http://github.com/deepmind/dm-haiku.
  32. 32.Hermann, K. L. and Lampinen, A. K. What shapes feature representations? exploring datasets, architectures, and training. In Advances in Neural Information Processing Systems, 2020.
  33. 33.Hermann, K. M., Hill, F., Green, S., Wang, F., Faulkner, R., Soyer, H., Szepesvari, D., Czarnecki, W. M., Jaderberg, M., Teplyashin, D., et al. Grounded language learning in a simulated 3d world. arXiv preprint arXiv:1706.06551, 2017.
  34. 34.Hill, F., Lampinen, A., Schneider, R., Clark, S., Botvinick, M., McClelland, J. L., and Santoro, A. Environmental drivers of systematicity and generalization in a situated agent. In International Conference on Learning Representations, 2019.
  35. 35.Holyoak, K. J. and Lu, H. Emergence of relational reasoning. Current Opinion in Behavioral Sciences, 37:118–124, 2021.
  36. 36.Humplik, J., Galashov, A., Hasenclever, L., Ortega, P. A., Teh, Y. W., and Heess, N. Meta reinforcement learning as task inference. arXiv preprint arXiv:1905.06424, 2019.
  37. 37.Ichien, N., Liu, Q., Fu, S., Holyoak, K. J., Yuille, A., and Lu, H. Visual analogy: Deep learning versus compositional models. In Proceedings of the 30th Annual Meeting of the Cognitive Science Society, 2021.
  38. 38.Jaderberg, M., Mnih, V., Czarnecki, W. M., Schaul, T., Leibo, J. Z., Silver, D., and Kavukcuoglu, K. Reinforcement learning with unsupervised auxiliary tasks. In International Conference on Learning Representations, 2016.
  39. 39.Jiang, Y., Gu, S., Murphy, K., and Finn, C. Language as an abstraction for hierarchical deep reinforcement learning. arXiv preprint arXiv:1906.07343, 2019.
  40. 40.Jumper, J., Evans, R., Pritzel, A., Green, T., Figurnov, M., Ronneberger, O., Tunyasuvunakool, K., Bates, R., Zˇ´ıdek, A., Potapenko, A., et al. Highly accurate protein structure prediction with alphafold. Nature, 596(7873):583–589, 2021.
  41. 41.Kaplan, R., Sauer, C., and Sosa, A. Beating atari with natural language guided reinforcement learning. arXiv preprint arXiv:1704.05539, 2017.
  42. 42.Katz, J. S. and Wright, A. A. Issues in the comparative cognition of same/different abstract-concept learning. Current Opinion in Behavioral Sciences, 37:29–34, 2021.
  43. 43.Keil, F. C., Wilson, R. A., and Wilson, R. A. Explanation and cognition. MIT press, 2000.
  44. 44.Kingma, D. P. and Ba, J. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014.
  45. 45.Kirk, R., Zhang, A., Grefenstette, E., and Rocktaschel, T. A survey of generalisation in deep reinforcement learning. arXiv preprint arXiv:2111.09794, 2021.
  46. 46.Lake, B. M., Ullman, T. D., Tenenbaum, J. B., and Gershman, S. J. Building machines that learn and think like people. Behavioral and brain sciences, 40, 2017.
  47. 47.Lertvittayakumjorn, P. and Toni, F. Explanation-based human debugging of nlp models: A survey. arXiv preprint arXiv:2104.15135, 2021.
  48. 48.Lombrozo, T. The structure and function of explanations. Trends in cognitive sciences, 10(10):464–470, 2006.
  49. 49.Lombrozo, T. and Carey, S. Functional explanation and the function of explanation. Cognition, 99(2):167–204, 2006.
  50. 50.Lombrozo, T. and Vasilyeva, N. Causal explanation. Oxford handbook of causal reasoning, pp. 415–432, 2017.
  51. 51.Luketina, J., Nardelli, N., Farquhar, G., Foerster, J., Andreas, J., Grefenstette, E., Whiteson, S., and Rocktaschel, T. A survey of reinforcement learning informed by natural language. arXiv preprint arXiv:1906.03926, 2019.
  52. 52.Lupyan, G. Taking symbols for granted? is the discontinuity between human and nonhuman minds the product of external symbol systems? Behavioral and Brain Sciences, 31(2):140–141, 2008.
  53. 53.Lupyan, G. The centrality of language in human cognition. Language Learning, 66(3):516–553, 2016.
  54. 54.Marcus, G. The next decade in ai: four steps towards robust artificial intelligence. arXiv preprint arXiv:2002.06177, 2020.
  55. 55.Mu, J., Liang, P., and Goodman, N. Shaping visual representations with language for few-shot classification. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pp. 4823–4830, Online, July 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.acl-main.436. URL https://aclanthology.org/2020.acl-main.436.
  56. 56.Mu, J., Zhong, V., Raileanu, R., Jiang, M., Goodman, N., Rocktaschel, T., and Grefenstette, E. Improving intrinsic exploration with language abstractions. arXiv preprint arXiv:2202.08938, 2022.
  57. 57.Nam, A. J. H. and McClelland, J. L. What underlies rapid learning and systematic generalization in humans. arXiv preprint arXiv:2102.02926, 2021.
  58. 58.Parisotto, E., Song, F., Rae, J., Pascanu, R., Gulcehre, C., Jayakumar, S., Jaderberg, M., Kaufman, R. L., Clark, A., Noury, S., et al. Stabilizing transformers for reinforcement learning. In International Conference on Machine Learning, pp. 7487–7498. PMLR, 2020.
  59. 59.Pearl, J. Theoretical impediments to machine learning with seven sparks from the causal revolution. arXiv preprint arXiv:1801.04016, 2018.
  60. 60.Pearl, J. The seven tools of causal inference, with reflections on machine learning. Communications of the ACM, 62 (3):54–60, 2019.
  61. 61.Penn, D. C., Holyoak, K. J., and Povinelli, D. J. Darwin’s mistake: Explaining the discontinuity between human and nonhuman minds. Behavioral and Brain Sciences, 31 (2):109–130, 2008.
  62. 62.Puebla, G. and Bowers, J. S. Can deep convolutional neural networks learn same-different relations? bioRxiv preprint, 2021.
  63. 63.Raileanu, R., Goldstein, M., Yarats, D., Kostrikov, I., and Fergus, R. Automatic data augmentation for generalization in deep reinforcement learning. arXiv preprint arXiv:2006.12862, 2020.
  64. 64.Rakelly, K., Zhou, A., Finn, C., Levine, S., and Quillen, D. Efficient off-policy meta-reinforcement learning via probabilistic context variables. In Chaudhuri, K. and Salakhutdinov, R. (eds.), Proceedings of the 36th International Conference on Machine Learning, volume 97 of Proceedings of Machine Learning Research, pp. 5331–5340. PMLR, 09–15 Jun 2019. URL https://proceedings.mlr.press/v97/rakelly19a.html.
  65. 65.Rezende, D. J., Danihelka, I., Papamakarios, G., Ke, N. R., Jiang, R., Weber, T., Gregor, K., Merzic, H., Viola, F., Wang, J., et al. Causally correct partial models for reinforcement learning. arXiv preprint arXiv:2002.02836, 2020.
  66. 66.Rittle-Johnson, B. Promoting transfer: Effects of self-explanation and direct instruction. Child development, 77 (1):1–15, 2006.
  67. 67.Ross, A. S., Hughes, M. C., and Doshi-Velez, F. Right for the right reasons: Training differentiable models by constraining their explanations. In Proceedings of the Twenty-Sixth International Joint Conference on Artificial Intelligence, IJCAI-17, pp. 2662–2670, 2017. doi: 10.24963/ijcai.2017/371. URL https://doi.org/10.24963/ijcai.2017/371.
  68. 68.Santoro, A., Raposo, D., Barrett, D. G., Malinowski, M., Pascanu, R., Battaglia, P., and Lillicrap, T. A simple neural network module for relational reasoning. arXiv preprint arXiv:1706.01427, 2017.
  69. 69.Santoro, A., Hill, F., Barrett, D., Morcos, A., and Lillicrap, T. Measuring abstract reasoning in neural networks. In International Conference on Machine Learning, pp. 4477–4486, 2018.
  70. 70.Santoro, A., Lampinen, A., Mathewson, K., Lillicrap, T., and Raposo, D. Symbolic behaviour in artificial intelligence. arXiv preprint arXiv:2102.03406, 2021.
  71. 71.Schramowski, P., Stammer, W., Teso, S., Brugger, A., Herbert, F., Shao, X., Luigs, H.-G., Mahlein, A.-K., and Kersting, K. Making deep neural networks right for the right scientific reasons by interacting with their explanations. Nature Machine Intelligence, 2(8):476–486, 2020.
  72. 72.Shanahan, M., Nikiforou, K., Creswell, A., Kaplanis, C., Barrett, D., and Garnelo, M. An explicitly relational neural network architecture. In International Conference on Machine Learning, pp. 8593–8603. PMLR, 2020.
  73. 73.Sinapov, J. and Stoytchev, A. The odd one out task: Toward an intelligence test for robots. In 2010 IEEE 9th International Conference on Development and Learning, pp. 126–131. IEEE, 2010.
  74. 74.Stammer, W., Schramowski, P., and Kersting, K. Right for the right concept: Revising neuro-symbolic concepts by interacting with their explanations. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 3619–3629, 2021.
  75. 75.Stephens, R. G. and Navarro, D. J. One of these greebles is not like the others: Semi-supervised models for similarity structures. Proceedings of the 30th Annual Meeting of the Cognitive Science Society, 2008.
  76. 76.Tam, A. C., Rabinowitz, N. C., Lampinen, A. K., Roy, N. A., Chan, S. C., Strouse, D., Wang, J. X., Banino, A., and Hill, F. Semantic exploration from language abstractions and pretrained representations. arXiv preprint arXiv:2204.05080, 2022.
  77. 77.Topin, N. and Veloso, M. Generation of policy-level explanations for reinforcement learning. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 33, pp. 2514–2521, 2019.
  78. 78.Tulli, S., Wallkotter, S., Paiva, A., Melo, F. S., and Chetouani, M. Learning from explanations and demonstrations: A pilot study. In 2nd Workshop on Interactive Natural Language Technology for Explainable Artificial Intelligence, pp. 61–66, 2020.
  79. 79.Van Fraassen, B. The pragmatic theory of explanation. Theories of Explanation, 8:135–155, 1988.
  80. 80.Wang, J. X., King, M., Porcel, N., Kurth-Nelson, Z., Zhu, T., Deck, C., Choy, P., Cassin, M., Reynolds, M., Song, F., et al. Alchemy: A structured task distribution for meta-reinforcement learning. arXiv preprint arXiv:2102.02926, 2021.
  81. 81.Watkins, O., Gupta, A., Darrell, T., Abbeel, P., and Andreas, J. Teachable reinforcement learning via advice distillation. Advances in Neural Information Processing Systems, 34, 2021.
  82. 82.Williams, J. J. and Lombrozo, T. The role of explanation in discovery and generalization: Evidence from category learning. Cognitive science, 34(5):776–806, 2010.
  83. 83.Wood, D., Bruner, J. S., and Ross, G. The role of tutoring in problem solving. Journal of child psychology and psychiatry, 17(2):89–100, 1976.
  84. 84.Woodward, J. Making things happen: A theory of causal explanation. Oxford university press, 2005.
  85. 85.Xie, N., Ras, G., van Gerven, M., and Doran, D. Explainable deep learning: A field guide for the uninitiated. arXiv preprint arXiv:2004.14545, 2020.

Citation

MLA
Lampinen, A. K., et al. “Tell Me Why! Explanations Support Learning Relational and Causal Structure”. International Conference on Machine Learning, vol. 162, 2022, pp. 11868–90, https://proceedings.mlr.press/v162/lampinen22a.html.
APA
Lampinen, A. K., Roy, N., Dasgupta, I., Chan, S. C., Tam, A., Mcclelland, J., Yan, C., Santoro, A., Rabinowitz, N. C., Wang, J., & Hill, F. (2022). Tell me why! Explanations support learning relational and causal structure. International Conference on Machine Learning, 162, 11868–11890. https://proceedings.mlr.press/v162/lampinen22a.html
Chicago
Lampinen, A. K., N. Roy, I. Dasgupta, et al. 2022. “Tell Me Why! Explanations Support Learning Relational and Causal Structure”. International Conference on Machine Learning 162: 11868–90. https://proceedings.mlr.press/v162/lampinen22a.html.
Harvard
Lampinen, A.K. et al. (2022) “Tell me why! Explanations support learning relational and causal structure”, International Conference on Machine Learning. PMLR, pp. 11868–11890. Available at: https://proceedings.mlr.press/v162/lampinen22a.html.
Vancouver
1. Lampinen AK, Roy N, Dasgupta I, et al (2022) Tell me why! Explanations support learning relational and causal structure. In: International Conference on Machine Learning. PMLR, pp 11868–11890

BibTeX

@InProceedings{pmlr-v162-lampinen22a,
  title = 	 {Tell me why! {E}xplanations support learning relational and causal structure},
  author =       {Lampinen, Andrew K and Roy, Nicholas and Dasgupta, Ishita and Chan, Stephanie Cy and Tam, Allison and Mcclelland, James and Yan, Chen and Santoro, Adam and Rabinowitz, Neil C and Wang, Jane and Hill, Felix},
  booktitle = 	 {Proceedings of the 39th International Conference on Machine Learning},
  pages = 	 {11868--11890},
  year = 	 {2022},
  editor = 	 {Chaudhuri, Kamalika and Jegelka, Stefanie and Song, Le and Szepesvari, Csaba and Niu, Gang and Sabato, Sivan},
  volume = 	 {162},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {17--23 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v162/lampinen22a/lampinen22a.pdf},
  url = 	 {https://proceedings.mlr.press/v162/lampinen22a.html},
  abstract = 	 {Inferring the abstract relational and causal structure of the world is a major challenge for reinforcement-learning (RL) agents. For humans, language{—}particularly in the form of explanations{—}plays a considerable role in overcoming this challenge. Here, we show that language can play a similar role for deep RL agents in complex environments. While agents typically struggle to acquire relational and causal knowledge, augmenting their experience by training them to predict language descriptions and explanations can overcome these limitations. We show that language can help agents learn challenging relational tasks, and examine which aspects of language contribute to its benefits. We then show that explanations can help agents to infer not only relational but also causal structure. Language can shape the way that agents to generalize out-of-distribution from ambiguous, causally-confounded training, and explanations even allow agents to learn to perform experimental interventions to identify causal relationships. Our results suggest that language description and explanation may be powerful tools for improving agent learning and generalization.}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/