What Does the Credential Still Certify? Cognitive Stewardship for AI-Mediated Education

Paper · arXiv 2607.19988 · Published July 22, 2026
Co-Writing and Collaboration

Generative AI is changing a basic premise of educational assessment: that submitted work can reliably evidence the human capacities a credential claims to certify. The challenge is not simply whether students use AI, but what remains inferable about learning when some cognitive work has been delegated to a system. This paper develops cognitive stewardship, a framework for AI-mediated assessment that links the learning claim, delegation boundary, evidence standard, and safeguards. We then audit verified public generative AI assessment guidance from 30 universities. Using a pre-specified scoring codebook–a written, source-grounded rubric–four open-weight LLM models applied the rubric as structured coders, with scores averaged to reduce dependence on any single model’s bias. The audit shows that public policies are becoming better at classifying AI use than at explaining what evidence and protections preserve credential validity. Boundaries are more visible than evidence standards; safeguards are uneven; and guidance is clearest when AI use resembles final-output substitution rather than feedback, access, verification, or professional workflow. The takeaway is that permission categories are necessary but insufficient.

Introduction. Generative AI has made a quiet premise of educational assessment newly fragile: that a submitted artifact can stand as evidence of a learner’s competence. A polished essay, program, proof, lesson plan, literature review, or design proposal may still show understanding. It may also reflect intensive machine assistance, private coaching, hidden outsourcing, or a legitimate accessibility support that is difficult to reconstruct after submission. The resulting problem is not only misconduct. It is whether the work handed in still supports the human claim a grade, course, or credential makes. This paper calls that problem educational delegation. The key question is not whether AI touched the work, but which cognitive operations moved from the learner to the system and which remained with the learner. One student may use AI feedback while retaining problem formulation, source evaluation, revision judgment, and final responsibility. Another may delegate topic selection, evidence search, argu- ment structure, drafting, citation, and prose revision.

Discussion / Conclusion. Generative AI changes what educational institutions can validly certify when learners may delegate parts of the work. The answer is not simply prohibition, permission, detection, or outsourcing. This paper has framed the problem as educational delegation and proposed cognitive stewardship: connect the learning claim, delegation boundary, evidence standard, and safeguard layer before treating a product as evidence of competence. The policy audit supports that diagnosis. The audited universities were not silent about generative AI; many had official pages, AI-use categories, and disclosure language. The gap was specific: rules about allowed use outpaced evidence for what credentials still certify. Boundary scores exceeded evidence scores for most policy packages, scenario guidance was clearest for final-output substitution, and safeguards appeared unevenly.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

Why do people disclose to AI systems despite their artificial nature? Why do some clarifying approaches produce understanding while others just satisfy? Why does polished presentation create unearned authority in AI outputs? Can local safety checks guarantee system-level behavioral safety? How can infrastructure records verify actual agent behavior? Does AI assistance promote real skill development or substitute for independent learning? What design and behavioral factors drive false consciousness attribution to AI? Do writers recognize when AI writing assistance alters their expressed stance? What linguistic features distinguish AI-generated text from human writing most reliably? How does the generation-verification gap limit what we can measure about AI reasoning? How should designers communicate what AI systems truly are and can do? How does AI adoption across firms reshape employment and inequality? How do false presuppositions and sycophancy drive persistent false beliefs in models?