Theme of inquiry
How do training signals and methods affect model learning and stability?
A question within its area, explored through 10 lines of inquiry below — each a family of specific questions the research asks.
57 specific questions
- Why does context information fail to override prior training associations?
- Does foundational model training or user priors more strongly shape final outputs?
- How does training order affect knowledge acquisition in language models?
- How do training associations override context information in language models?
- Why does consistency training make models resistant to prompt perturbations?
- Can data filtering during pretraining prevent cognitive biases in language models?
- How do training-data priors influence model defaults when context is ambiguous?
54 specific questions
- Why does self-correction during generation produce reliable labels without exemplars?
- How does error avalanching compound failures in self-training iterations?
- How does error distribution during training affect a model's ability to self-correct?
- Can models learn to generate their own training examples effectively?
- Can self-consistency checks fully prevent error avalanching in self-training loops?
- What failure modes emerge when model-generated content trains on itself iteratively?
- Why does self-consistency fail as a proxy reward for correctness?
70 specific questions
- Why does latent-level prediction beat token-level prediction for reasoning?
- Why do language models generate reasoning tokens after internally deciding the answer?
- How does reward density during training affect token efficiency in reasoning?
- Why does token-level gradient targeting matter more than aggregate loss?
- Can models internally identify which tokens matter most for reasoning?
- How do meta-tokens help models learn when to generate reasoning versus commit predictions?
- Does the token prediction framing actually capture what human reasoning does?
53 specific questions
- Can structured natural language feedback outperform scalar rewards in RL?
- Why does alternating RL training stabilize learning better than simultaneous updates?
- Can decomposing consistency into multiple metrics improve reinforcement learning for dialogue?
- Can emotion-grounded rewards replace coarse bonus signals in hierarchical dialogue RL?
- Can outcome-based rewards fully replace per-step likelihood in diffusion RL training?
- Can multi-turn reinforcement learning improve tool use in language models?
- Does semantic diversity in output space compete with reward-component diversity?
45 specific questions
- Why do longer sequences tolerate higher sparsity than shorter ones?
- Should production deployments scale budgets with sequence length for sparse models?
- Do task-relevant parameter changes naturally concentrate in sparse regions?
- Does static per-token sparsity repeat the fixed-budget mistake at short sequences?
- What makes sparse attention more reliable for long-context retrieval?
- Why does representation sparsity reliably indicate task difficulty for language models?
- How does task type interact with sequence length in sparsity tolerance?
99 specific questions
- Can RL format selection explain performance gains attributed to algorithmic improvements?
- How do surface statistical regularities enable correct outputs while degrading robustness?
- Can diversity-aware RL objectives prevent format convergence?
- Does parameter isolation per task enable online updates without retraining?
- When does natural context diversity reduce the need for explicit exploration?
- Why do parameter-based compressors fail to measure true model simplicity?
- What makes output convergence across models inevitable given input-side homogenization?
80 specific questions
- How does scaling and training data enable compositional behavior without symbolic mechanisms?
- Does the linear representation hypothesis reflect networks or reflect our analysis tools?
- Does latent density emerge during pretraining from training data familiarity?
- How does representational density emerge from training data familiarity?
- Can fractured representations explain why models fail at systematic generalization?
- Why does gradient descent discover compositional structure without explicit pressure?
- Can neural networks represent symbolic structures without explicit mechanisms?
74 specific questions
- Can recurrent transformers learn genuinely new computations beyond inference stages?
- Can looping enable reasoning capabilities that fixed-depth transformers fundamentally cannot achieve?
- Why does looping computation outperform adding more transformer layers?
- Can bounded-depth transformers solve inherently sequential problems?
- Can transformers reason beyond fixed architectural depth limits?
- Why do standard transformers fail to encode recursive structure in their hidden states?
- Can latent recurrence achieve the depth that standard transformers cannot?
67 specific questions
- Can smaller models actually perform well on specific downstream tasks?
- How does model scale affect the crossover point between base and post-trained performance?
- Does scaling model size solve compositional generalization problems?
- Why do larger models reduce interference between rare and common tasks?
- Does scaling data automatically produce compositional reasoning or just better feature encoding?
- How do task frequency and complexity interact with model capacity during training?
- Why does exploration quality matter more than learner network depth?
86 specific questions
- Does teacher-style refinement of training data transfer equally to all student model distributions?
- At what point does output quality outweigh diversity value in synthetic data tasks?
- How do quality and diversity in synthetic data interact with accumulation schedules?
- Can selecting the right data subset outperform training on everything?
- How does the ratio of synthetic to real training data affect model collapse?
- How do quality, diversity, and complexity create different effects on downstream model performance?
- Why does diversity of training cases matter more than raw dataset size?