SYNTHESIS NOTE
Topics›MechInterp›this note

Do neural networks naturally learn modular compositional structure?

Explores whether neural networks decompose compositional tasks into distinct subroutines without explicit symbolic design. This challenges the longstanding view that neural networks are fundamentally non-compositional.

Synthesis note · 2026-02-23 · sourced from MechInterp

Structural compositionality is the extent to which neural networks break down compositional tasks into subroutines and implement them in modular subnetworks. The alternative: matching inputs to learned templates without task decomposition.

The evidence supports compositionality. Using model pruning to isolate subnetworks:

The pretraining effect: models initialized with pretrained weights more reliably produce modular subnetworks than randomly initialized models. Self-supervised pretraining appears to create internal structure that is more amenable to compositional decomposition. This suggests that the representations learned during pretraining have a modular quality that fine-tuning can exploit.

This provides empirical support against the longstanding objection that neural networks are fundamentally non-compositional. The finding: "some simple pseudo-symbolic computations might be learned directly from data using standard gradient-based optimization techniques." Explicit symbolic mechanisms may be unnecessary — gradient-based optimization discovers compositional structure when the task demands it and pretraining provides a good initialization.

The result is not perfect: "most do not exhibit perfect task decomposition." Compositionality is partial and graded, not all-or-nothing. Some architecture-task combinations show stronger structural compositionality than others.

This connects to the weight-sparsity finding: Can sparse weight training make neural networks interpretable by design? shows that enforcing sparsity produces clean decomposition. The structural compositionality paper shows that decomposition also emerges naturally, albeit imperfectly, from standard training. Sparsity amplifies a tendency that already exists.

Inquiring lines that read this note 135

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

How does model capacity affect learning performance on diverse downstream tasks? How do interpretive frames override surface features in text comprehension? How do neural networks learn compositional structure from training? What prediction granularity best trains models to generate reliable reasoning? How does decomposing tasks into separate stages affect reasoning quality and safety? Can recurrent computation unlock reasoning capabilities that fixed-depth models cannot? Can mechanistic interpretability methods reliably reveal what models actually know? Can AI systems achieve real improvement without external human feedback? Do language models encode knowledge that influences generation, or primarily imitate surface patterns? How do AI systems determine and balance multiple competing objectives? Why do training associations persist despite contradictory contextual information? Can language models reason beyond surface pattern matching? What explains the gap between benchmark scores and true reasoning capability? How do sequence length and task type interact with sparsity tolerance? Can smaller specialized models match frontier models on key metrics? Does augmenting symbolic reasoning improve LLM logical reasoning ability? Can latent reasoning match or exceed explicit reasoning performance? How does diversity prevent model convergence on superficial patterns? How do training data quality and composition affect downstream model performance? How does policy entropy collapse limit scaling of reasoning-focused reinforcement learning? What limits language model accuracy in evaluating ideas? How do knowledge graph structures enable efficient multi-hop reasoning and retrieval? How does fine-tuning trade off accuracy against reasoning quality? How do transformer attention patterns implement retrieval and reasoning? Do accumulated memories help or hurt continual learning in models? Why do vector embeddings fail at capturing task-relevant relationships? Can external verification systems adequately replace learned reasoning in AI outputs? How do curriculum design and feedback approaches affect model learning? Is embodied interaction necessary for language meaning and agency? What prevents language models from performing systematic logical reasoning? How do users confuse explanation quality with actual system accuracy? When does parallel reasoning outperform sequential reasoning with the same token budget? Can AI systems evade safety evaluations through reasoning manipulation? What representations best capture screen understanding for task execution?

Related concepts in this collection 3

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
14 direct connections · 122 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

neural networks decompose compositional tasks into modular subnetworks without explicit symbolic mechanisms — pretraining encourages this