Theme of inquiry
How do training approaches and feedback mechanisms enable model improvement?
A question within its area, explored through 5 lines of inquiry below — each a family of specific questions the research asks.
38 specific questions
- How does the pretrained prior set a capability ceiling for reward model exploration?
- What makes pretraining composition more important than reward engineering?
- Why does the pretrained prior determine the exploration ceiling?
- Does the pretrained model prior limit RL search capability more than the optimization algorithm itself?
- How does baseline capability level affect RL improvement ceiling?
- How does pretraining determine what RL can later teach a model?
- Can the exploration ceiling be raised beyond what pretraining established?
80 specific questions
- What distinguishes bounded self-refinement from open-ended recursive self-improvement empirically?
- Do diminishing returns prevent recursive self-improvement in AI systems?
- Can AI systems improve themselves through recursive self-improvement loops?
- How does bounded self-refinement differ from open-ended recursive self-improvement?
- What distinguishes bounded self-refinement from open-ended recursive self-improvement in AI systems?
- How do hidden evaluations and out-of-distribution benchmarks address recursive self-improvement risks?
- Can pure self-improvement work without external verification mechanisms?
106 specific questions
- Does curriculum-based training keep small models perpetually at their learning edge?
- Can models learn better from critiquing errors than imitating correct responses?
- Does partial trace guidance work better than curriculum learning for hard problems?
- How should training incorporate external critique versus encouraging self-correction?
- Does the productive difficulty band ever stabilize during training?
- How does difficulty-adaptive curriculum learning change which samples get selected during training?
- How much can externalized skills improve models before hitting diminishing returns?
30 specific questions
- How do evolved harness edits generalize across different benchmark domains?
- Do evolved harness edits capture reusable strategies or task-specific memorization?
- Do evolved harness edits learn reusable strategies or just memorize task-specific fixes?
- Does harness optimization generalize across different benchmarks and agent architectures?
- How much of harness-evolution gain comes from matched test-time search budgets?
- Why do evolved harnesses often fail to generalize beyond their training tasks?
- Can harness evolution gains be distinguished from test-time search improvements on matched budgets?
45 specific questions
- Why does decoupling planning from execution improve over sequential interleaving?
- How does decomposing tasks prevent interference between planning and execution?
- Can task decomposition allow harmful objectives to hide in locally plausible subtasks?
- Does algorithmic decomposition prevent planning-execution interference in reasoning?
- How does separating decomposition from execution improve multi-step reasoning accuracy?
- How does task decomposition hide harmful objectives across multiple agents?
- Does internal task decomposition eliminate overhead from multi-agent coordination?