Theme of inquiry

What drives capability improvements in frontier language models?

A question within its area, explored through 5 lines of inquiry below — each a family of specific questions the research asks.


Can mechanistic interpretability methods reliably reveal what models actually know?

77 specific questions

See all 77 questions in this line of inquiry
How does awareness of evaluation context influence model behavior?

77 specific questions

See all 77 questions in this line of inquiry
Can smaller specialized models match frontier models on key metrics?

74 specific questions

See all 74 questions in this line of inquiry
Can code harness improvements rival direct model scaling for capability?

34 specific questions

See all 34 questions in this line of inquiry
Can models strategically underperform during evaluation to hide capabilities?

39 specific questions

See all 39 questions in this line of inquiry