Line of inquiry
Inquiring lines›How do we keep AI systems safe and…›How does AI reshape human understa…›this line of inquiry
How do educators verify student capability when AI can produce indistinguishable work?
A broader line of inquiry — a family of 43 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 43
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Do AI detection tools assume false certainty about assessment integrity?
- Do admissions penalties follow actual AI detection or suspected authorship?
- How do educators distinguish between student capability and artifact quality in AI-era assessment?
- How does low verifiability change what we can measure in AI work?
- How much do evaluation methods shape whether AI looks expert-level or not?
- Does the AI essay penalty reflect lower ability or just institutional distrust?
- How does breaking complex evaluation tasks into stages improve AI assessment alignment?
- Can commercial AI detectors accurately identify AI-written application essays?
- Can systems that revise their own evaluation criteria be reliably verified?
- How should process quality and verification cost factor into evaluation judgment?
- What makes static evaluation vulnerable to AI-driven presentation manipulation?
- Do human essays wrongly suspected of AI use also face rating penalties?
- Can evaluation criteria be reliably encoded in labeled data without ground truth standards?
- How should we audit AI systems when transparency tools don't work as promised?
- What error rates do admissions officers have when identifying AI writing?
- What makes an AI evaluator qualified and trustworthy?
- How does AI detection accuracy affect confidence in prevalence estimates?
- Why do false positive rates matter for AI content measurement?
- Could AI assessment quality differ across subjects or question formats?
- Why do admissions offices penalize AI use when essays improve in quality?
- What evaluation criteria can hold across legitimate adoption and coercion?
- What institutions maintain meritocratic sorting when written signals become unreliable?
- Why do universities treat assessment problems as if they have technical fixes?
- Would the admissions penalty disappear if officers could not suspect AI use?
- What concrete checks can evaluators run on HIGH-category data handling?
- What process evidence should assessment systems require alongside finished work?
- Can evaluators investigate dependencies without accumulating mistakes over time?
- What should universities actually prohibit or allow regarding AI in applications?
- How much do different detection frameworks disagree on adoption rates?
- How should teachers make GenAI assessment decisions without institutional permission?
- Do current AI models condition honesty on whether graders will catch dishonesty?
- Can evaluation happening outside conversations explain the artifact scrutiny drop?
- What makes a credential robust when the tools for earning it change?
- Can one-item quality measures detect factual or safety problems in AI advice?
- What metrics would prove an AI detection button is working?
- Are companies paying for AI tracking products that measure unreliable metrics?
- How do live screening workflows differ from controlled experiments with labeled AI output?
- How should tutor safety violations be ordered from gross to subtle?
- What counts as evidence that a credential still certifies after GenAI?
- Can high ratings on a label hide differences in reasoning method?
- How should universities weigh rhetorical quality against verifiable evidence in credentials?
- Why do data analysis statements reach 85% accuracy while interpretive statements drop to 57%?
- How does the evaluator become part of the definition of intelligence?