Line of inquiry
Inquiring lines›How do we keep AI systems safe and…›How does AI reshape human understa…›this line of inquiry
How do real-world evaluations reveal AI capabilities that benchmarks hide?
A broader line of inquiry — a family of 50 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 50
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Why do open-world evaluations reveal capabilities that static benchmarks hide?
- How do live human evaluations differ from ground-truth benchmarks?
- Why do AI benchmarks measure accuracy instead of reasoning quality?
- How do existing evaluations measure AI capability in contained environments?
- Does higher capability correlate with more benchmark contamination?
- Do multi-axis benchmarks reveal failures that single-axis benchmarks systematically hide?
- What gap exists between AI model capability in benchmarks and real client work?
- How do open-world evaluations correct distortions that automated benchmarks introduce?
- How can high benchmark performance mask broken reasoning in AI systems?
- How much do internal benchmarks differ from independent model evaluations?
- How do benchmark environments misrepresent deployment readiness?
- Do standard benchmarks miss how humans actually fail to use AI advice?
- Do narrow benchmark improvements translate directly to broader economic capability gains?
- Can optimization metrics hide actual versus apparent progress in AI systems?
- How should domain-specific AI be evaluated differently from general benchmarks?
- Can automated benchmarks accurately capture progress on real-world long-horizon tasks?
- Should evaluations shift toward open-world messy tasks instead of contests?
- Why do benchmark tasks differ from real occupational workflows in practice?
- Do automated benchmarks accurately measure real-world strategic reasoning ability?
- Which separable capability axes reveal when single benchmarks misrepresent deployment readiness?
- Why do static benchmarks miss frontier capabilities that open-world tasks reveal?
- Why do benchmark scores not capture the true nature of AI systems?
- How do surface correlations between narratives and answers mislead benchmark validity?
- How should human-AI evaluation differ from standalone model benchmarks?
- Can self-administered surveys establish trustworthy AI capability benchmarks?
- Can open-world evaluations become a scalable paradigm without becoming the next benchmark trap?
- How do lab-scale benchmark tasks differ from real frontier AI research?
- How do narrow benchmark optimizations differ from genuine architectural research discoveries?
- Why does benchmark saturation give a false sense of capability coverage?
- Can verified test performance substitute for subjective judgment about capability?
- Do frontier AI models fail in ways that preserve the appearance of competence?
- Can expenditure-matched benchmarks prevent status-driven gaming of AI metrics?
- How do single average metrics conceal rare but severe AI failures?
- How should evaluation frameworks account for the computational cost of frontier AI capability?
- Why do interactive tasks show larger capability gaps than bounded tasks?
- What distinct domains of AI competence do current assessments actually measure?
- Can a single benchmark score capture both progress and readiness?
- Can self-reported AI reliability metrics hide confounding factors like task complexity?
- Can frontier AI models match expert human performance on specialized tasks?
- Why do isolated model benchmarks understate real-world AI security risks?
- What drives the gap between AI capability and actual cost savings in practice?
- Can capability claims be fact-checked when labs control the process narrative?
- Why do AI benchmarks show rapid saturation from near-zero to near-perfect?
- How do traditional quality assurance methods fail for mutable AI outputs?
- What makes synthetic control benchmarks representative of real deployment misalignment risks?
- Can expert-frontier exams discriminate frontier capability better than saturated benchmarks?
- How does capability evaluation differ from alignment evaluation in difficulty?
- What capability gap prevents GenAI systems from moving beyond pilots?
- What capability dimensions does a single aggregate pass rate hide?
- What would whole-system AGI evaluation look like in practice?