Can a plain AI model catch specialized designs just by training on better data, or does its architecture set limits?
Can dense models match specialized architectures by mixing data better?
This explores whether a standard dense model, where every parameter is used for every input, can close the gap with specialized designs like Mixture-of-Experts or memory-augmented models just by training on a better blend of data, rather than by changing its architecture.
This explores whether a plain dense model can catch up with specialized architectures through smarter data mixing alone. The collection has no paper that tests this head-on, with data recipes on one side and architectures on the other. What it does have are several findings that point in one direction: data and training shape what a model becomes, but architecture sets limits that data alone doesn't seem to remove.
The clearest evidence against "data is enough" comes from work on how a model's capacity is laid out. Engram pairs simple N-gram lookup memory with Mixture-of-Experts routing. At the same parameter count and the same compute, the hybrid beats pure MoE, and the best results come from a balanced split between the two mechanisms (Can lookup memory and computation work together better than either alone?). Surprisingly, the gains show up most in reasoning and code, not in simple fact lookup. This suggests some capabilities come from how memory and computation are wired together, which is not something you can mix in through data. At the small end of the scale, the shape of a dense model matters too. Deep-and-thin models beat balanced ones by several points at 125M–350M parameters (Does depth matter more than width for tiny language models?). So even within "dense," architecture choices aren't neutral.
There's a more interesting angle: dense models may already contain specialized structure. Pruning experiments show that networks split compositional tasks into isolated subnetworks without being told to, and pretraining makes this modular structure more consistent (Do neural networks naturally learn modular compositional structure?). In other words, a dense model may grow informal "experts" inside itself, and the data it sees helps decide how cleanly they form. A caution comes with this: matching scores can hide very different internal organization. A model can reach perfect accuracy while its internal structure is fragmented and breaks under distribution shift (Can models be smart without organized internal structure?). A dense model that "matches" a specialized one on a benchmark may not match it in robustness.
The training-regime findings show what data and training can and can't do. Reasoning-trained models stay ahead of non-reasoning models no matter how much inference compute the latter get. How a model was trained matters more than raw budget (Can non-reasoning models catch up with more compute?). Small models can match large ones on function calling when trained with DPO on the teacher's correct and incorrect examples (Can small models match large models on function calling?). But that win is narrow and task-specific, not a general closing of the gap. There's also a warning about mixing itself. RL post-training tends to pick one format from the pretraining mix within the first epoch and suppress the rest (Does RL training collapse format diversity in pretrained models?). A carefully balanced data mix may not stay balanced once later training stages act on it.
The pattern across these notes: "matching" usually needs something beyond more or better data. That might be a structural addition like lookup memory, a training signal that teaches a protocol, or an external check. For example, committees of weak model calls only match strong models when tests or proofs can pick out the right answer (When can weak models match strong model performance?). If you want the direct data-mixing experiments, such as DoReMi-style mixture optimization compared against MoE, the collection doesn't cover them yet.
Sources 8 notes
Engram combines O(1) N-gram lookup with Mixture-of-Experts routing, revealing a U-shaped scaling law where balanced allocation to both mechanisms outperforms either alone. Gains appear largest in reasoning and code rather than pure retrieval.
MobileLLM shows deep-and-thin architectures yield 2.7–4.3% accuracy gains over balanced designs at 125M–350M scale by composing abstract concepts through layers rather than spreading parameters across width.
Pruning experiments reveal that neural networks implement compositional subroutines in isolated subnetworks, with ablations affecting only their corresponding function. Pretraining substantially increases the consistency and reliability of this modular structure across architectures and domains.
Models trained with SGD can contain all the linearly decodable features needed for a task while maintaining fundamentally broken internal organization. This makes them vulnerable to perturbation and distribution shift invisible to standard evaluation metrics.
Reasoning models persistently outperform non-reasoning models regardless of inference budget because training instills a reasoning protocol that makes additional tokens productive. The gap is fundamentally about deployment mechanisms and training structure, not raw capability.
Show all 8 sources
Small models fine-tuned via DPO on correct and incorrect function-calling examples from a large teacher model achieve high accuracy on logical and mathematical tasks. DPO's explicit negative examples directly target the rigid output format failures where SFT alone underperforms.
Controlled experiments show RL consistently amplifies one format distribution from pretraining within the first epoch while collapsing alternatives. The winning format depends on model scale, not necessarily performance, and is largely hidden when starting from proprietary pretrained models.
Sampling alone amplifies coverage but cannot select correct solutions. Reliable performance matching requires external soundness signals—tests, proofs, or type checks—that convert latent correct proposals into actual selections.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Break It Down: Evidence for Structural Compositionality in Neural Networks
- Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining
- Hierarchical Reasoning Model
- MobileLLM: Optimizing Sub-billion Parameter Language Models for On-Device Use Cases
- Single-Agent LLMs Outperform Multi-Agent Systems on Multi-Hop Reasoning Under Equal Thinking Token Budgets
- Sharpening Tax in Post-Training
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learning
- A Survey on Test-Time Scaling in Large Language Models: What, How, Where, and How Well?