Line of inquiry
Inquiring lines›What drives capability improvement…›How do training signals and method…›this line of inquiry
How do neural networks learn compositional structure from training?
A broader line of inquiry — a family of 80 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 80
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- How does scaling and training data enable compositional behavior without symbolic mechanisms?
- Does the linear representation hypothesis reflect networks or reflect our analysis tools?
- Does latent density emerge during pretraining from training data familiarity?
- How does representational density emerge from training data familiarity?
- Can fractured representations explain why models fail at systematic generalization?
- Why does gradient descent discover compositional structure without explicit pressure?
- Can neural networks represent symbolic structures without explicit mechanisms?
- Can scaling alone create compositional generalization without explicit binding mechanisms?
- Where do neural networks still fail at compositional generalization despite scaling?
- How do neural networks decompose tasks into modular subnetworks that transfer?
- What inductive bias would force models to learn Newtonian mechanics instead of shortcuts?
- Can sparse approximations reveal interpretable structure hidden in existing dense models?
- Does representational density emerge from training data exposure during pretraining?
- Can we detect and measure circuit formation before generalization emerges?
- What role does a model's representational structure play in learning?
- Can steering vectors prove that representations are genuinely organized?
- What are fractured entangled representations in neural networks?
- Can training order and structure shape what networks retain and learn?
- How do concept vectors in neural networks relate to human-language chess explanations?
- How do latents at the same hierarchy level become more correlated than tokens?
- How do semantic features in representations become steerable task-specific directions?
- Why does weight sparsity reduce superposition and force disentangled representations?
- Can representation engineering cleanly isolate single features in entangled semantic space?
- Can fractured entangled representations hide undetected by standard analysis methods?
- How does weight sharing compound the advantages of deeper model designs?
- Can world models form from aggregated partial information across training distributions?
- What makes a new representational primitive valuable enough to justify its representational cost?
- What makes a feature abstract versus concrete in neural network activations?
- How do models develop dense representations for familiar training data?
- What inductive biases help networks segregate entities from raw inputs?
- How should we rethink the symbolism versus connectionism debate in light of LLMs?
- Why does knowledge storage separate from reasoning circuits in neural networks?
- What happens to representational structure during model pretraining phases?
- How much do structural inductive biases matter compared to training data volume?
- Do KANs maintain their advantages in deep architectures and large-scale training?
- Can neural networks implement genuine algorithms or only statistical pattern matching?
- How do neural networks decompose complex tasks into modular subnetworks?
- How do cortical columns implement local inference over memory cycles?
- Can we predict which tasks will decompose into modular subnetworks?
- Does architectural discovery follow an empirical scaling law like neural networks?
- How can neural networks be interpretable by design rather than post-hoc?
- What makes linear decodability a reliable signal of compositionality?
- What prevents representation collapse in latent-prediction world models like JEPA?
- What role does query-level exposure play in enabling compositional generalization?
- How do knowledge and reasoning circuits interfere in the same neural network?
- Why does adaptation concentrate in low-dimensional subspaces of weights or representations?
- Does the same spectral signature appear across different embedding models?
- How do sparse circuits compare to the modular subnetworks that emerge naturally?
- Can universal function approximators be expensive to learn in practice?
- Could probing methods miss computationally important features in neural networks?
- What neural or architectural mechanism allows selective override of frequency effects?
- How does representation-level reranking address residual gaps after decomposition?
- Can generative reconstruction preserve latent manifold structure better than geometric compression?
- What limits the extrapolation of learned operations like rotation and reflection?
- How do encode-decode contractive biases create stable attractors in latent space?
- What makes regularization an implicit factor in embedding geometry?
- What solvable idealized settings reveal fundamental phenomena in realistic deep learning?
- How do overparameterization and data size shift what attractors represent?
- What makes multimodal conditioning effective when features are decomposed to the right granularity?
- Can curvature measurements predict task difficulty without behavioral labels?
- Can autoencoders act as associative memory systems like Hopfield networks?
- What happens when a single loss function conflates representation learning with decision-making?
- Do feature extraction methods systematically miss computationally important complex features?
- How do classical mechanics and statistical mechanics provide methodological templates for learning theory?
- How does joint backpropagation differ from training separate ensemble models?
- Can gradient approximation at equilibrium replace backpropagation through time in practice?
- What non-parametric methods could replace latent factors for inductive learning?
- How do weight visualizations reveal temporal structure in cyclic training?
- What physical structure does a Gaussian-regularized latent space actually encode?
- Why are polysemantic features concentrated in early neural network layers?
- How do gradients flowing through both branches simultaneously reshape each component's role?
- Which hyperparameter theories best explain universal behaviors across neural networks?
- How do biological brains organize computation across different cortical timescales?
- Why does projecting lattice vectors to a subfield increase shared unit distances?
- How do embedding dimension limits constrain what concept models can represent?
- How do neural networks extend contextual bandits beyond linear reward assumptions?
- How does latent space diffusion enable evolutionary search in high dimensions?
- What makes data augmentation an implicit form of contraction learning?
- Why do different brain and AI systems appear similar when compared via RSA?
- Do substitute networks converge differently than complement networks?