Break It Down: Evidence for Structural Compositionality in Neural Networks
Michael A. LeporiThomas SerreEllie Pavlick
Demonstrates through network pruning that neural models implicitly decompose complex vision and language tasks into modular, independently functional subnetworks, challenging the assumption that compositionality requires explicit symbolic mechanisms.
Modern artificial intelligence systems achieve strong performance across computer vision and language processing, yet how they organize internal computations remains heavily debated. A central question is whether these models learn compositionality—breaking complex tasks into distinct, reusable subroutines and combining them—or merely match inputs against holistic memorized templates. Understanding this distinction is vital for determining whether building reliable AI requires integrating specialized symbolic architectures or if standard deep learning models can naturally acquire modular reasoning directly from training data.
The article evaluates whether standard neural networks exhibit structural compositionality by implicitly implementing subroutines within isolated, modular subnetworks. To investigate this, the authors used model pruning techniques, specifically continuous sparsification, to identify whether dedicated parameter subsets compute specific subroutines within trained vision and language models. They evaluated vision models (ResNet-50 and Wide ResNet-50) on visual reasoning tasks involving spatial relations and object counts, alongside language models (BERT-Small) on syntactic agreement tasks, testing both models trained from scratch and those initialized with self-supervised pretraining.
The findings provide substantial evidence that neural networks frequently learn modular task decomposition without explicit symbolic mechanisms. First, the authors isolated specialized subnetworks that performed their target subroutines with high accuracy while failing on alternative subroutines. Second, when these discovered subnetworks were ablated (removed) from the base models, accuracy on the target subroutine dropped sharply while performance on non-targeted subroutines remained largely intact, demonstrating genuine internal modularity. Third, self-supervised pretraining significantly enhanced this modular organization in language models, yielding more consistent and robust subnetwork separation compared to training from scratch. Finally, transformer vision models (ViT) failed to achieve baseline accuracy on these visual reasoning tasks, while control experiments confirmed that these modular structures do not appear in randomly initialized networks.
These results carry significant strategic implications for artificial intelligence development and safety. By demonstrating that neural networks can autonomously acquire modular, pseudo-symbolic representations, the findings suggest organizations may not need to overhaul deep learning pipelines with complex, costly symbolic add-ons to achieve structured reasoning. Furthermore, discovering discrete functional subnetworks enhances mechanistic interpretability, opening clearer pathways to audit, control, and edit specific internal model behaviors to mitigate operational and compliance risks.
Moving forward, technical teams should leverage subnetwork pruning and ablation methods to audit internal task modularity and utilize pretraining to foster cleaner functional representations. However, decision-makers should exercise measured confidence: the current methodology requires pre-defining target subroutines, carries significant computational search costs (requiring thousands of GPU-hours), and was evaluated on tightly controlled synthetic benchmarks. Future research and pilot implementations must clarify the link between internal structural modularity and real-world out-of-distribution generalization, while developing automated, unsupervised techniques to discover unknown subroutines in complex, open-ended operational settings.
- Paper: Is a Modular Architecture Enough?, Sarthak Mittal et al. (2022). Provides a direct diagnostic benchmark on whether modular architectures naturally specialize into functional units, framing the core question of structural modularity and compositionality.
- Paper: MoEfication: Transformer Feed-forward Layers are Mixtures of Experts, Zhengyan Zhang et al. (2022). Demonstrates how feed-forward layers in standard transformers naturally partition into modular sub-networks of specialized experts, providing foundational evidence for structural modularity.
- Paper: Network Dissection: Quantifying Interpretability of Deep Visual Representations, David Bau et al. (2017). Introduces quantitative methods for isolating and dissecting concept-specific units in deep representations, which underpins the study's approach to identifying modular subnetworks.
- Paper: Compositional Generalization in Unsupervised Compositional Representation Learning: A Study on Disentanglement and Emergent Language, Zhenlin Xu et al. (2022). Investigates compositional generalization across learned representations, framing the key debate on whether explicit inductive biases or emergent neural structures enable compositionality.
- Paper: Rethinking the Value of Network Pruning, Zhuang Liu et al. (2019). Establishes foundational insights into the behavior of structured and unstructured model pruning, which the source utilizes as its primary investigative tool.
- Paper: Object Detectors Emerge in Deep Scene CNNs, Bolei Zhou et al. (2014). Shows that discrete functional subroutines emerge spontaneously inside deep neural networks without explicit component supervision.
- Paper: Optimal Brain Damage, Yann LeCun et al. (1989). Introduces saliency-based network ablation and parameter pruning to isolate and remove non-essential subnetworks.
- Paper: Towards Automated Circuit Discovery for Mechanistic Interpretability, Arthur Conmy et al. (2023). Extends the discovery of modular subnetworks by automating the identification of precise computational circuits underlying specific behaviors using activation patching.
- Paper: Mechanistically analyzing the effects of fine-tuning on procedurally defined tasks, Samyak Jain et al. (2024). Applies subnetwork pruning and mechanistic analysis to investigate how downstream adaptation affects localized, modular capabilities learned during pretraining.
- Paper: How Transformers Learn Causal Structure with Gradient Descent, Eshaan Nichani et al. (2024). Provides formal mathematical and mechanistic foundations explaining how gradient descent drives transformers to learn modular causal subroutines and induction circuits.
- Paper: Hilbert Operator for Progressive Encoding (HOPE): A Mathematical Framework for Deconstructing Learned Representations in Deep Networks, Hossein Mobahi et al. (2026). Develops a functional operator-space framework for capacity reduction and representation deconstruction, generalizing the pruning-based subnetwork analysis.
- Paper: DepGraph: Towards Any Structural Pruning, Gongfan Fang et al. (2023). Builds upon structural parameter interdependence across diverse architectures to automate structural dependency grouping and subnetwork pruning.
