Break It Down: Evidence for Structural Compositionality in Neural Networks

Michael A. LeporiThomas SerreEllie Pavlick

article2023NeurIPS66 citations

Demonstrates through network pruning that neural models implicitly decompose complex vision and language tasks into modular, independently functional subnetworks, challenging the assumption that compositionality requires explicit symbolic mechanisms.

Listen

Modern artificial intelligence systems achieve strong performance across computer vision and language processing, yet how they organize internal computations remains heavily debated. A central question is whether these models learn compositionality—breaking complex tasks into distinct, reusable subroutines and combining them—or merely match inputs against holistic memorized templates. Understanding this distinction is vital for determining whether building reliable AI requires integrating specialized symbolic architectures or if standard deep learning models can naturally acquire modular reasoning directly from training data.

The article evaluates whether standard neural networks exhibit structural compositionality by implicitly implementing subroutines within isolated, modular subnetworks. To investigate this, the authors used model pruning techniques, specifically continuous sparsification, to identify whether dedicated parameter subsets compute specific subroutines within trained vision and language models. They evaluated vision models (ResNet-50 and Wide ResNet-50) on visual reasoning tasks involving spatial relations and object counts, alongside language models (BERT-Small) on syntactic agreement tasks, testing both models trained from scratch and those initialized with self-supervised pretraining.

The findings provide substantial evidence that neural networks frequently learn modular task decomposition without explicit symbolic mechanisms. First, the authors isolated specialized subnetworks that performed their target subroutines with high accuracy while failing on alternative subroutines. Second, when these discovered subnetworks were ablated (removed) from the base models, accuracy on the target subroutine dropped sharply while performance on non-targeted subroutines remained largely intact, demonstrating genuine internal modularity. Third, self-supervised pretraining significantly enhanced this modular organization in language models, yielding more consistent and robust subnetwork separation compared to training from scratch. Finally, transformer vision models (ViT) failed to achieve baseline accuracy on these visual reasoning tasks, while control experiments confirmed that these modular structures do not appear in randomly initialized networks.

These results carry significant strategic implications for artificial intelligence development and safety. By demonstrating that neural networks can autonomously acquire modular, pseudo-symbolic representations, the findings suggest organizations may not need to overhaul deep learning pipelines with complex, costly symbolic add-ons to achieve structured reasoning. Furthermore, discovering discrete functional subnetworks enhances mechanistic interpretability, opening clearer pathways to audit, control, and edit specific internal model behaviors to mitigate operational and compliance risks.

Moving forward, technical teams should leverage subnetwork pruning and ablation methods to audit internal task modularity and utilize pretraining to foster cleaner functional representations. However, decision-makers should exercise measured confidence: the current methodology requires pre-defining target subroutines, carries significant computational search costs (requiring thousands of GPU-hours), and was evaluated on tightly controlled synthetic benchmarks. Future research and pilot implementations must clarify the link between internal structural modularity and real-world out-of-distribution generalization, while developing automated, unsupervised techniques to discover unknown subroutines in complex, open-ended operational settings.

arXiv: 2301.10884
Cover for Break It Down: Evidence for Structural Compositionality in Neural Networks

Abstract

Though modern neural networks have achieved impressive performance in both vision and language tasks, we know little about the functions that they implement. One possibility is that neural networks implicitly break down complex tasks into subroutines, implement modular solutions to these subroutines, and compose them into an overall solution to a task — a property we term structural compositionality. Another possibility is that they may simply learn to match new inputs to learned templates, eliding task decomposition entirely. Here, we leverage model pruning techniques to investigate this question in both vision and language across a variety of architectures, tasks, and pretraining regimens. Our results demonstrate that models often implement solutions to subroutines via modular subnetworks, which can be ablated while maintaining the functionality of other subnetworks. This suggests that neural networks may be able to learn compositionality, obviating the need for specialized symbolic mechanisms.

Table of Contents

  • 1 Introduction
  • 2 Structural Compositionality
  • 3 Experimental Design
  • 3.1 Preliminaries
  • 3.2 Experimental Logic
  • 4 Discovering Subnetworks
  • 5 Vision Experiments
  • 6 Language Experiments
  • 7 Results
  • 8 Effect of Pretraining on Structural Compositionality
  • 9 Related Work
  • 10 Discussion
  • Acknowledgments and Disclosure of Funding
  • References
  • A Full Results
  • A.1 Statistical Analysis of Main Results
  • B Continuous Sparsification: Extended Discussion
  • C Mask Hyperparameter Search Details
  • D ViT Hyperparameter Search Results
  • E Vision Stimuli
  • F Language Data Details
  • G Vision Pretraining Details
  • H Subnetwork Sparsity Data
  • I Control Experiment: Random Models
  • I.1 Statistical Analysis of Main Results vs. Random Results
  • J Pruned Model Analysis
  • K Subnetwork Overlap Analysis

Knowls

  1. Knowl 1 — Structural compositionality as modular subroutine implementation

    definition

    Structural compositionality is the extent to which a neural network decomposes a task into binary subroutines and implements those subroutines in modular parts of its parameters, rather than encoding only their conjunction in an entangled template. A compositional rule combines two subroutines as C(x)=SR1(x)∧SR2(x)C(x)=SR_1(x)\land SR_2(x), where each SRiSR_i maps an input xx to a binary output.

    For a base model MCM_C with weight parameters θ\theta, a subnetwork for subroutine SRiSR_i is defined by a binary weight mask mi∈{0,1}dm_i\in\{0,1\}^d:

    Sub⁡i(x)=MC(x;θ⊙mi),\operatorname{Sub}_i(x)=M_C(x;\theta\odot m_i),

    where dd is the number of masked weight parameters and ⊙\odot denotes elementwise multiplication. Bias parameters are not masked. The corresponding ablated model retains the complementary parameters, using θ⊙(1−mi)\theta\odot(1-m_i). Structural compositionality therefore requires both functional specialization of subnetworks and modularity under ablation; perfect decomposition of the entire task is not required.

  2. Knowl 2 — Subnetwork and ablation test for structural compositionality

    model/method

    The method first trains a base model on an odd-one-out task whose rule is the conjunction of two binary subroutines. For each target subroutine SRiSR_i, a binary mask is then learned so that the masked model solves an odd-one-out dataset requiring SRiSR_i. The masked model is evaluated on two matched datasets: Test Target Subroutine contains examples whose odd-one-out can be identified by computing SRiSR_i, whereas Test Other Subroutine contains examples requiring the other subroutine.

    The same two evaluations are performed after ablating the learned subnetwork. Let atarget(M)a_{\mathrm{target}}(M) and aother(M)a_{\mathrm{other}}(M) denote accuracy on these datasets, with all accuracies clamped to [0.25,1][0.25,1] because four-way odd-one-out classification has chance accuracy 0.250.25. The diagnostic differences are

    Δsub=atarget(Sub⁡i)−aother(Sub⁡i),\Delta_{\mathrm{sub}}=a_{\mathrm{target}}(\operatorname{Sub}_i)-a_{\mathrm{other}}(\operatorname{Sub}_i), Δabl=atarget(Mablated,i)−aother(Mablated,i).\Delta_{\mathrm{abl}}=a_{\mathrm{target}}(M_{\mathrm{ablated},i})-a_{\mathrm{other}}(M_{\mathrm{ablated},i}).

    A modular subnetwork should yield Δsub>0\Delta_{\mathrm{sub}}>0, because it specializes in its target subroutine, while its ablation should yield Δabl<0\Delta_{\mathrm{abl}}<0, because the target computation is impaired more than the other computation. An entangled, non-compositional solution is expected to produce approximately equal performance on the two datasets before and after ablation.

  3. Knowl 3 — Continuous sparsification for discovering subnetworks

    algorithm

    For a frozen base model with weights θ\theta, the mask for subroutine SRiSR_i is optimized to preserve performance on the subroutine task while minimizing the number of active weights. The ideal binary objective is

    min⁡mi∈{0,1}d  LSRi ⁣(MC(⋅;θ⊙mi))+λ∥mi∥1,\min_{m_i\in\{0,1\}^d}\;L_{SR_i}\!\left(M_C(\cdot;\theta\odot m_i)\right)+\lambda\lVert m_i\rVert_1,

    where mim_i is a binary mask over dd weight parameters, LSRiL_{SR_i} is the odd-one-out loss for the target subroutine, and λ\lambda controls sparsity. Continuous sparsification replaces mim_i with σ(βsi)\sigma(\beta s_i), where si∈Rds_i\in\mathbb{R}^d is optimized, σ\sigma is the elementwise sigmoid, and β\beta is an increasing temperature:

    min⁡si∈Rd  LSRi ⁣(MC(⋅;θ⊙σ(βsi)))+λ∥σ(βsi)∥1.\min_{s_i\in\mathbb{R}^d}\;L_{SR_i}\!\left(M_C(\cdot;\theta\odot\sigma(\beta s_i))\right)+\lambda\lVert\sigma(\beta s_i)\rVert_1.

    The temperature is increased exponentially toward βmax⁡=200\beta_{\max}=200, so the soft mask approaches the Heaviside mask H(s)H(s), with H(s)=0H(s)=0 for s<0s<0 and H(s)=1H(s)=1 for s>0s>0. The model weights remain frozen; only mask parameters are trained. The reported mask optimization uses Adam, batch size 64, 90 epochs, λ=10−8\lambda=10^{-8}, learning-rate candidates {0.01,0.0001}\{0.01,0.0001\}, and mask initializations {0.1,0.05,0,−0.05}\{0.1,0.05,0,-0.05\}. The search also varies where masking begins: the whole network, the final two stages or layers, or only the final stage or layer, depending on architecture. A candidate is retained only if it reaches at least 90% accuracy on its target subroutine task, after which the configuration maximizing the ablated-model compositionality score is selected.

  4. Knowl 4 — Controlled vision tasks and model architectures

    experimental setup

    The vision experiments construct odd-one-out datasets from three binary image subroutines: whether one shape is inside another, whether the shapes are in contact, and whether an image contains a specified number of shapes. The resulting compositional rules are Inside-Contact, Number-Contact, and Inside-Number. For Inside-Contact, the four image types are (+Inside,+Contact)(+\mathrm{Inside},+\mathrm{Contact}), (−Inside,+Contact)(-\mathrm{Inside},+\mathrm{Contact}), (+Inside,−Contact)(+\mathrm{Inside},-\mathrm{Contact}), and (−Inside,−Contact)(-\mathrm{Inside},-\mathrm{Contact}). In the base task, only (+Inside,+Contact)(+\mathrm{Inside},+\mathrm{Contact}) follows the rule; in each four-image example, three images follow the rule and one is the odd-one-out. The analogous Number tasks use NN shapes for +Number+\mathrm{Number} and a different count M≠NM\ne N for −Number-\mathrm{Number}.

    Each vision model consists of a backbone followed by a two-layer MLP with hidden size 2048 and output size 128. The model embeds all four images, compares their embeddings by pairwise dot products, and predicts the least similar image as the odd-one-out using cross-entropy loss. The backbones are ResNet50, Wide ResNet50, and Vision Transformer; BatchNorm is replaced with InstanceNorm. Three random seeds are trained for each condition using Adam, early stopping for at most 100 epochs with patience 75, without dropout or weight decay. Vision Transformer models failed to exceed 90% accuracy on any compositional task and were therefore excluded from the structural analysis.

  5. Knowl 5 — Controlled language tasks and BERT architecture

    experimental setup

    The language experiments apply the same odd-one-out and masking procedure to two syntactic agreement domains. Subject-Verb Agreement requires matching the syntactic number of a subject and a distant verb; Reflexive Anaphora requires matching the number of an antecedent and a reflexive pronoun. Each example contains four sentences, three rule-following sentences and one rule-breaking sentence.

    For Subject-Verb Agreement, the compositional rule pairs a singular subject with a singular verb or a plural subject with a plural verb. For Reflexive Anaphora, the corresponding rule pairs a singular antecedent with a singular pronoun or a plural antecedent with a plural pronoun. The datasets are split into singular and plural partitions so that rule-following examples are grammatical and rule-breaking examples are ungrammatical. Separate subroutine datasets target the subject versus verb number or antecedent versus pronoun number, and matched evaluation datasets isolate the target and other subroutines.

    The language model is BERT-Small with four hidden layers. The sentence representation is the embedding of its [CLS] token, and unlike the vision models, it does not use an additional MLP head.

  6. Knowl 6 — Neural networks frequently contain modular solutions to task subroutines

    empirical result

    Across the trained ResNet50, Wide ResNet50, and BERT-Small models that solved their base compositional tasks, the discovered subnetworks generally showed the predicted positive target-minus-other accuracy difference, while their ablated complements generally showed the predicted negative difference. This pattern occurred without an explicit modularity or compositionality objective. Generalized linear-model analyses found subnetwork differences significantly different from zero across the reported architecture–task–subroutine combinations, and ablation differences were significant in many combinations despite the small number of trained models.

    The result was not uniformly strong for every subroutine. In ResNet50 models trained on Number-Contact, the Contact computation was localized in a relatively small, modular subnetwork, whereas the Number computation was more diffusely represented and produced weaker evidence. Layerwise overlap measurements also indicated that subnetworks for the same subroutine were more reproducible than subnetworks for different subroutines. Thus, the models often exhibit structural compositionality, but usually not complete or perfectly factorized task decomposition.

  7. Knowl 7 — Pretraining improves compositional modularity in language models

    empirical result

    Initializing BERT-Small with pretrained language-model weights made modular subnetwork structure more reliable across the Subject-Verb Agreement and Reflexive Anaphora tasks than training BERT-Small from random initialization. The pretrained models more consistently produced a strong disparity between performance on the target and other subroutine datasets, and the mask-discovery procedure appeared less sensitive to random seed.

    The vision comparison used ResNet50 weights obtained by SimCLR self-supervised pretraining on the compositional image data, with temperature 0.070.07, batch size 256, 100 pretraining epochs, and a learning rate starting at 0.00050.0005 under cosine decay. Unlike the language results, the vision experiments showed no obvious improvement from pretraining over random initialization. The paper therefore supports a pretraining benefit for structural compositionality in language, but not for the tested vision setting.

  8. Knowl 8 — Random-network controls separate learned modularity from mask expressivity

    empirical result

    The authors repeated the mask-learning procedure on randomly initialized ResNet50 and BERT-Small networks using the same mask hyperparameters as for trained models. Random networks could sometimes yield subnetworks with higher accuracy on Test Target Subroutine than on Test Other Subroutine, showing that target-specialized behavior alone can arise from mask optimization.

    However, their ablation behavior differed from that of trained models. In random ResNet50 networks, ablating the discovered masks generally reduced performance to chance on both evaluation datasets rather than selectively damaging the target subroutine. In random BERT-Small networks, ablation often produced the opposite of the compositional pattern: performance remained higher on the target dataset than on the other-subroutine dataset. These controls indicate that the trained-model results depend on causal structure in learned representations rather than only on the expressive power of binary masks.

  9. Knowl 9 — Subnetworks for different subroutines have low parameter overlap

    empirical result

    For one ResNet50 trained on the Inside-Number task, the authors discovered three masks for the Inside subroutine and three masks for the Number subroutine, then measured per-layer intersection-over-union (IoU), where for two parameter sets AA and BB,

    IoU⁡(A,B)=∣A∩B∣∣A∪B∣.\operatorname{IoU}(A,B)=\frac{|A\cap B|}{|A\cup B|}.

    Within the Inside subroutine, IoU values across the masked layers ranged from 0.9600.960 to 0.9940.994, indicating highly reproducible masks. Within the Number subroutine, values ranged from 0.1290.129 to 0.5160.516, indicating noisier localization. The IoU between the intersections of the three Inside masks and the three Number masks ranged from 0.0270.027 to 0.3210.321, consistently below the within-subroutine overlap and especially low in earlier backbone layers. This supports the claim that distinct subroutines occupy largely different parameter subsets, although the localization can be subroutine-dependent.

  10. Knowl 10 — Methodological limitations and unresolved relationship to generalization

    limitation

    The mask-discovery method requires the researcher to specify the candidate subroutines in advance; it does not discover unknown functional decompositions autonomously. Its interpretation also depends on causal ablation tests and control models, because a performant masked network is not by itself evidence that the corresponding parameters are causally responsible for the base model's behavior. In addition, the relationship between structural compositionality—modular organization inside the weights—and compositional generalization to new combinations remains unresolved. The mask hyperparameter search is computationally expensive because each model–subroutine pair requires training many separate masks.

Coverage note — The supplementary pruned-model replication, exhaustive sparsity tables, and illustrative stimulus and sentence examples were omitted because they mainly corroborate the main masking procedure or instantiate the task definitions rather than add distinct load-bearing findings.

References

  1. 1.Andreas, J., Rohrbach, M., Darrell, T., and Klein, D. Neural module networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 39–48, 2016.
  2. 2.Bau, A., Belinkov, Y., Sajjad, H., Durrani, N., Dalvi, F., and Glass, J. Identifying and controlling important neurons in neural machine translation. In International Conference on Learning Representations, 2019. URL https://openreview.net/forum?id=H1z-PsR5KX.
  3. 3.Bhargava, P., Drozd, A., and Rogers, A. Generalization in nli: Ways (not) to go beyond simple heuristics. In Proceedings of the Second Workshop on Insights from Negative Results in NLP, pp. 125–135, 2021.
  4. 4.Black, S., Sharkey, L., Grinsztajn, L., Winsor, E., Braun, D., Merizian, J., Parker, K., Guevara, C. R., Millidge, B., Alfour, G., et al. Interpreting neural networks through the polytope lens. arXiv preprint arXiv:2211.12312, 2022.
  5. 5.Cammarata, N., Carter, S., Goh, G., Olah, C., Petrov, M., Schubert, L., Voss, C., Egan, B., and Lim, S. K. Thread: Circuits. Distill, 5(3):e24, 2020.
  6. 6.Cao, S., Sanh, V., and Rush, A. M. Low-complexity probing via finding subnetworks. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pp. 960–966, 2021.
  7. 7.Chen, T., Kornblith, S., Norouzi, M., and Hinton, G. A simple framework for contrastive learning of visual representations. In International conference on machine learning, pp. 1597–1607. PMLR, 2020.
  8. 8.Chughtai, B., Chan, L., and Nanda, N. A toy model of universality: Reverse engineering how networks learn group operations. arXiv preprint arXiv:2302.03025, 2023.
  9. 9.Csordás, R., van Steenkiste, S., and Schmidhuber, J. Are neural nets modular? inspecting functional modularity through differentiable weight masks. In International Conference on Learning Representations, 2021.
  10. 10.Dankers, V., Bruni, E., and Hupkes, D. The paradox of the compositionality of natural language: A neural machine translation case study. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 4154–4175, 2022.
  11. 11.De Cao, N., Schlichtkrull, M. S., Aziz, W., and Titov, I. How do decisions emerge across layers in neural models? interpretation with differentiable masking. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp. 3243–3255, 2020.
  12. 12.De Cao, N., Schmid, L., Hupkes, D., and Titov, I. Sparse interventions in language models with differentiable masking. In Proceedings of the Fifth BlackboxNLP Workshop on Analyzing and Interpreting Neural Networks for NLP, pp. 16–27, 2022.
  13. 13.Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018.
  14. 14.Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al. An image is worth 16x16 words: Transformers for image recognition at scale. In International Conference on Learning Representations, 2020.
  15. 15.Dziri, N., Lu, X., Sclar, M., Li, X. L., Jian, L., Lin, B. Y., West, P., Bhagavatula, C., Bras, R. L., Hwang, J. D., et al. Faith and fate: Limits of transformers on compositionality. arXiv preprint arXiv:2305.18654, 2023.
  16. 16.Elazar, Y., Ravfogel, S., Jacovi, A., and Goldberg, Y. Amnesic probing: Behavioral explanation with amnesic counterfactuals. Transactions of the Association for Computational Linguistics, 9:160–175, 2021.
  17. 17.Ellis, K., Wong, L., Nye, M., Sable-Meyer, M., Cary, L., Anaya Pozo, L., Hewitt, L., Solar-Lezama, A., and Tenenbaum, J. B. Dreamcoder: growing generalizable, interpretable knowledge with wake–sleep bayesian program learning. Philosophical Transactions of the Royal Society A, 381(2251):20220050, 2023.
  18. 18.Ettinger, A., Elgohary, A., Phillips, C., and Resnik, P. Assessing composition in sentence vector representations. In Proceedings of the 27th International Conference on Computational Linguistics, pp. 1790–1801, 2018.
  19. 19.Fodor, J. A. and Lepore, E. The compositionality papers. Oxford University Press, 2002.
  20. 20.Ganguli, D., Hatfield-Dodds, Z., Hernandez, D., Jones, A., Kernion, J., Lovitt, L., Ndousse, K., Amodei, D., Brown, T., Clark, J., Kaplan, J., McCandlish, S., and Olan, C. A mathematical framework for transformer circuits. Transformer Circuits Thread, 2021.
  21. 21.Guo, D., Rush, A. M., and Kim, Y. Parameter-efficient transfer learning with diff pruning. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pp. 4884–4896, 2021.
  22. 22.He, K., Zhang, X., Ren, S., and Sun, J. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 770–778, 2016.
  23. 23.Henighan, T., Carter, S., Humne, T., Elhage, N., Lasenby, R., Fort, S., Schiefer, N., and Olah, C. Superposition, memorization, and double descent. Transformer Circuits Thread, 2023.
  24. 24.Hod, S., Filan, D., Casper, S., Critch, A., and Russell, S. Quantifying local specialization in deep neural networks. arXiv e-prints, pp. arXiv–2110, 2022.
  25. 25.Hupkes, D., Dankers, V., Mul, M., and Bruni, E. Compositionality decomposed: How do neural networks generalise? Journal of Artificial Intelligence Research, 67:757–795, 2020.
  26. 26.Kim, N. and Linzen, T. Cogs: A compositional generalization challenge based on semantic interpretation. In Empirical Methods in Natural Language Processing, 2020.
  27. 27.Kim, N., Linzen, T., and Smolensky, P. Uncontrolled lexical exposure leads to overestimation of compositional generalization in pretrained models. arXiv preprint arXiv:2212.10769, 2022.
  28. 28.Kingma, D. P. and Ba, J. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014.
  29. 29.Kirby, S., Cornish, H., and Smith, K. Cumulative cultural evolution in the laboratory: An experimental approach to the origins of structure in human language. Proceedings of the National Academy of Sciences, 105(31):10681–10686, 2008.
  30. 30.Koh, P. W., Nguyen, T., Tang, Y. S., Mussmann, S., Pierson, E., Kim, B., and Liang, P. Concept bottleneck models. In International Conference on Machine Learning, pp. 5338–5348. PMLR, 2020.
  31. 31.Lake, B. and Baroni, M. Generalization without systematicity: On the compositional skills of sequence-to-sequence recurrent networks. In International conference on machine learning, pp. 2873–2882. PMLR, 2018.
  32. 32.Lake, B. M. Compositional generalization through meta sequence-to-sequence learning. Advances in neural information processing systems, 32, 2019.
  33. 33.Lake, B. M., Ullman, T. D., Tenenbaum, J. B., and Gershman, S. J. Building machines that learn and think like people. Behavioral and brain sciences, 40, 2017.
  34. 34.Lippe, P. Tutorial 17: Self-supervised contrastive learning with simclr. https://github.com/phlippe/uvadlc_notebooks/blob/master/docs/tutorial_notebooks/tutorial17/SimCLR.ipynb, 2022.
  35. 35.Louizos, C., Welling, M., and Kingma, D. P. Learning sparse neural networks through l_0 regularization. In International Conference on Learning Representations, 2018.
  36. 36.Lovering, C. and Pavlick, E. Unit testing for concepts in neural networks. Transactions of the Association for Computational Linguistics, 10:1193–1208, 2022.
  37. 37.Mandelbaum, E., Dunham, Y., Feiman, R., Firestone, C., Green, E., Harris, D., Kibbe, M. M., Kurdi, B., Mylopoulos, M., Shepherd, J., et al. Problems and mysteries of the many languages of thought. Cognitive Science, 46(12):e13225, 2022.
  38. 38.Marcus, G. F. The algebraic mind: Integrating connectionism and cognitive science. MIT press, 2003.
  39. 39.Marvin, R. and Linzen, T. Targeted syntactic evaluation of language models. Proceedings of the Society for Computation in Linguistics (SCiL), pp. 373–374, 2019.
  40. 40.Merrill, W., Tsilivis, N., and Shukla, A. A tale of two circuits: Grokking as competition of sparse and dense subnetworks. In ICLR 2023 Workshop on Mathematical and Empirical Understanding of Foundation Models, 2023.
  41. 41.Mhaskar, H. N. and Poggio, T. Deep vs. shallow networks: An approximation theory perspective. Analysis and Applications, 14(06):829–848, 2016.
  42. 42.Nye, M., Solar-Lezama, A., Tenenbaum, J., and Lake, B. M. Learning compositional rules via neural program synthesis. Advances in Neural Information Processing Systems, 33:10832–10842, 2020.
  43. 43.Olah, C. Mechanistic interpretability, variables, and the importance of interpretable bases. Transformer Circuits Thread, 2022.
  44. 44.Piantadosi, S. T., Tenenbaum, J. B., and Goodman, N. D. The logical primitives of thought: Empirical foundations for compositional cognitive models. Psychological review, 123(4):392, 2016.
  45. 45.Poggio, T. How deep sparse networks avoid the curse of dimensionality: Efficiently computable functions are compositionally sparse. CBMM Memo, 10:2022, 2022.
  46. 46.Qiu, L., Shaw, P., Pasupat, P., Nowak, P. K., Linzen, T., Sha, F., and Toutanova, K. Improving compositional generalization with latent structure and data augmentation. arXiv preprint arXiv:2112.07610, 2021.
  47. 47.Quilty-Dunn, J., Porot, N., and Mandelbaum, E. The best game in town: The re-emergence of the language of thought hypothesis across the cognitive sciences. Behavioral and Brain Sciences, pp. 1–55, 2022.
  48. 48.Ramanujan, V., Wortsman, M., Kembhavi, A., Farhadi, A., and Rastegari, M. What’s hidden in a randomly weighted neural network? In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 11893–11902, 2020.
  49. 49.Ravfogel, S., Elazar, Y., Gonen, H., Twiton, M., and Goldberg, Y. Null it out: Guarding protected attributes by iterative nullspace projection. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pp. 7237–7256, 2020.
  50. 50.Ren, Y., Guo, S., Labeau, M., Cohen, S. B., and Kirby, S. Compositional languages emerge in a neural iterated learning model. In International Conference on Learning Representations, 2019.
  51. 51.Savarese, P., Silva, H., and Maire, M. Winning the lottery with continuous sparsification. Advances in Neural Information Processing Systems, 33:11380–11390, 2020.
  52. 52.Saxton, D., Grefenstette, E., Hill, F., and Kohli, P. Analysing mathematical reasoning abilities of neural models. In International Conference on Learning Representations, 2018.
  53. 53.Smolensky, P., McCoy, R., Fernandez, R., Goldrick, M., and Gao, J. Neurocompositional computing: From the central paradox of cognition to a new generation of ai systems. AI Magazine, 43(3):308–322, 2022.
  54. 54.Tucker, M., Qian, P., and Levy, R. What if this modified that? syntactic interventions with counterfactual embeddings. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pp. 862–875, 2021.
  55. 55.Turc, I., Chang, M., Lee, K., and Toutanova, K. Well-read students learn better: The impact of student initialization on knowledge distillation. CoRR, abs/1908.08962, 2019. URL http://arxiv.org/abs/1908.08962.
  56. 56.Vani, A., Schwarzer, M., Lu, Y., Dhekane, E., and Courville, A. Iterated learning for emergent systematicity in vqa. In International Conference on Learning Representations, 2020.
  57. 57.Wortsman, M., Ramanujan, V., Liu, R., Kembhavi, A., Rastegari, M., Yosinski, J., and Farhadi, A. Supermasks in superposition. Advances in Neural Information Processing Systems, 33:15173–15184, 2020.
  58. 58.Wu, Z., Geiger, A., Rozner, J., Kreiss, E., Lu, H., Icard, T., Potts, C., and Goodman, N. Causal distillation for language models. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pp. 4288–4295, 2022.
  59. 59.Yu, L. and Ettinger, A. Assessing phrasal representation and composition in transformers. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp. 4896–4907, 2020.
  60. 60.Zagoruyko, S. and Komodakis, N. Wide residual networks. In British Machine Vision Conference 2016. British Machine Vision Association, 2016.
  61. 61.Zerroug, A., Vaishnav, M., Colin, J., Musslick, S., and Serre, T. A benchmark for compositional visual reasoning. In Thirty-sixth Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2022.
  62. 62.Zhang, D., Ahuja, K., Xu, Y., Wang, Y., and Courville, A. Can subnetwork structure be the key to out-of-distribution generalization? In International Conference on Machine Learning, pp. 12356–12367. PMLR, 2021.
  63. 63.Zhou, H., Lan, J., Liu, R., and Yosinski, J. Deconstructing lottery tickets: Zeros, signs, and the supermask. Advances in neural information processing systems, 32, 2019.
  64. 64.Zhu, W., Shaw, P., Linzen, T., and Sha, F. Learning to generalize compositionally by transferring across semantic parsing tasks. arXiv preprint arXiv:2111.05013, 2021.

Citation

MLA
Lepori, M., et al. “Break It Down: Evidence for Structural Compositionality in Neural Networks”. Advances in Neural Information Processing Systems, vol. 36, 2023, pp. 42623–60, https://proceedings.neurips.cc/paper_files/paper/2023/file/85069585133c4c168c865e65d72e9775-Paper-Conference.pdf.
APA
Lepori, M., Serre, T., & Pavlick, E. (2023). Break It Down: Evidence for Structural Compositionality in Neural Networks. Advances in Neural Information Processing Systems, 36, 42623–42660. https://proceedings.neurips.cc/paper_files/paper/2023/file/85069585133c4c168c865e65d72e9775-Paper-Conference.pdf
Chicago
Lepori, M., T. Serre, and E. Pavlick. 2023. “Break It Down: Evidence for Structural Compositionality in Neural Networks”. Advances in Neural Information Processing Systems 36: 42623–60. https://proceedings.neurips.cc/paper_files/paper/2023/file/85069585133c4c168c865e65d72e9775-Paper-Conference.pdf.
Harvard
Lepori, M., Serre, T. and Pavlick, E. (2023) “Break It Down: Evidence for Structural Compositionality in Neural Networks”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 42623–42660. Available at: https://proceedings.neurips.cc/paper_files/paper/2023/file/85069585133c4c168c865e65d72e9775-Paper-Conference.pdf.
Vancouver
1. Lepori M, Serre T, Pavlick E (2023) Break It Down: Evidence for Structural Compositionality in Neural Networks. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 42623–42660

BibTeX

@inproceedings{lepori2023break,
  title = {Break It Down: Evidence for Structural Compositionality in Neural Networks},
  author = {Lepori, Michael and Serre, Thomas and Pavlick, Ellie},
  year = {2023},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {36},
  pages = {42623-42660},
  url = {https://proceedings.neurips.cc/paper_files/paper/2023/file/85069585133c4c168c865e65d72e9775-Paper-Conference.pdf}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors