Theme of inquiry
How should computational architecture and inference strategies be tailored to different problem types?
A question within its area, explored through 5 lines of inquiry below — each a family of specific questions the research asks.
90 specific questions
- Why do benchmark scores rise while reasoning quality declines?
- Do reasoning benchmarks predict real performance in long delegated workflows?
- Can benchmark scores on verifiable tasks transfer to unseen problems outside the training domain?
- Why do estimates for task-level performance differ so much from full job automation timelines?
- How does measurement error in capability benchmarks systematically underestimate or overestimate true ability?
- Can identical model performance mask fundamentally broken internal representations?
- What is the gap between benchmark performance and real workplace task completion?
79 specific questions
- What mechanisms drive test-time compute allocation in reasoning tasks?
- Does test-time compute actually substitute for having larger model parameters?
- Can test-time compute allocation shift from solutions to strategies?
- How should we allocate compute between reasoning and retrieval iterations?
- Where does inference compute stop substituting for model capacity?
- Can inference budgets be allocated adaptively based on prompt difficulty?
- How should inference budgets adapt based on prompt difficulty?
69 specific questions
- Can context compression preserve what matters without introducing bias?
- Does including full context always degrade memory retrieval quality in practice?
- How does externalized state affect the long-context bottleneck in language models?
- How should memory consolidation timing differ across multiple timescales?
- Can compressed long-term memory outperform fixed-window token retention?
- What persistent memory architectures best support storing precomputed inferences across sessions?
- How do recurrent memory systems handle ultra-long context differently than attention?
52 specific questions
- Can parallel reasoning chains outperform longer sequential chains with the same compute?
- Can parallel thinking outperform sequential thinking under the same token budget?
- What makes parallel thinking more efficient than sequential chains?
- Why does parallel thinking outperform sequential thinking under fixed token budgets?
- Can parallel independent reasoning outperform sequential iterative refinement?
- When does sequential chain-of-thought dramatically beat parallel voting approaches?
- What advantages emerge from running 13 times more parallel reasoning chains with the same budget?
34 specific questions
- Can multiple small models outperform a single large model with good routing?
- What makes routing a better investment than training larger models?
- Should model routing decisions account for prompt-tier dependencies?
- How do routing and test-time compute scaling work together as optimization axes?
- Does model selection matter more than model improvement for query routing?
- Can embedding-cluster routing outperform a single frontier model?
- How does routing decide between models before generation happens?