Research Article
AI-Assisted Assessment and Instruction in Higher Education: Foundations, Applications, and Implications for Exam Design
Synthesis: Klapproth's preprint argues that large language models capable of completing conventional academic tasks at or above passing thresholds make the assumption behind much text-based Assessment untenable, and that the response should be principled redesign rather than detection and restriction. It moves in three stages. The foundations set out what LLMs are and what they cannot be relied on for: probabilistic continuation engines that hallucinate, shift with prompt wording, lack embodied understanding of causality, and vary systematically by task type, complexity and domain. The applications review how educators can use the same systems in their own workflow — structured prompting, retrieval-augmented generation, agentic item construction, automated scoring and formative Feedback. The exam-design implications join the two through constructive alignment: assessment should target the competences LLMs perform worst at. The practical upshot is a shift from AI-resistant to AI-robust design, formats that stay valid when AI is available because genuine student competence is required to complete them — practical and experimental tasks, oral and interactive formats, contextually situated assignments, and process documentation — with AI used openly for item generation, scoring support and feedback, and final grading decisions left to a qualified human examiner. The paper states plainly that it is a preprint, not peer reviewed, and discloses that a model was used for language editing only.
Key Findings
- Constructive alignment is the paper's organising frame. Following Biggs and Tang (2011), Assessment must be consistent with intended learning outcomes and teaching activities; a generative AI system that can produce a well-structured academic text indistinguishably from a student breaks that consistency, so the problem is one of design rather than detection.
- LLM performance falls as cognitive level rises. Huber and Niklaus (2025) mapped 43 Benchmark tasks from the technical reports of leading models onto Bloom's revised taxonomy: GPT-4 scored 0.99 at Remember, 0.89 at Understand, 0.93 at Apply but 0.76 at Analyze; Llama 3 fell from 0.90 to 0.59 across the same range; Claude 3 Haiku from 0.46 to 0.48 at much lower levels.
- The higher levels are unmeasured, not merely weak. No standard benchmark in that analysis contained tasks at the Evaluate or Create levels, and metacognitive knowledge was absent entirely, so what LLMs can do at higher-order thinking remains empirically uncharted even as those are the levels the paper argues higher education should assess.
- Traditional exam formats sit exactly where AI is strongest. Remembering, understanding and applying are both the levels current models handle most reliably and the levels conventional text-based examinations most commonly target, making those formats substantially susceptible to AI-assisted completion.
- Professional licensing exams are already passable. Mihalache et al. (2024) found recent chatbots reaching passing thresholds on all three stages of the USMLE, with errors concentrated on highly complex, interdisciplinary or specialized items; passable is not equal to reliable.
- Interaction quality degrades real-world performance. Bean et al. (2026) found LLMs identifying at least one relevant condition in 94.9% of standardized diagnostic cases, a figure that dropped substantially when real users rather than researchers posed the queries, because users supplied incomplete information and often failed to extract correct information present in the response.
- The design goal is robustness, not resistance. An AI-robust assessment validly measures intended outcomes even when students have AI access, not because AI cannot help but because genuine competence is necessary to complete it — moving the question from how to prevent AI use to which competences the task should require.
- Four format families carry that robustness. Practical and experimental tasks requiring embodied competence; oral and interactive formats such as viva voce examinations, presentations with examiner questioning and role plays; contextually situated and personalized tasks anchored in the student's own data, placement or institution; and process-oriented Assessment through portfolios, mandated drafts and reflective journals.
- Educator-side AI use is real but bounded. Structured prompting with the P.R.O.M.P.T. framework, retrieval-augmented generation for content validity, low temperature settings, a five-stage agentic item-construction workflow and a five-step scoring framework all appear as practical guidance, alongside limitations: scoring is a black box, LLMs score short undeveloped essays higher and long essays with minor errors lower than human raters do (Mathew et al., 2026), and final summative grading remains the responsibility of a human examiner.
What the evidence says AI can and cannot be relied on for
The technical foundation is stated briskly. LLMs are transformer-based systems trained on next-token prediction over very large corpora: they produce a reasonable continuation of the text so far, calibrated against patterns observed across billions of pages, which the paper takes to mean they are not reasoning systems in any classical sense. The capabilities it credits are substantial — analyzing and generating complex academic texts, multi-step problem solving over several hundred pages of context, interpreting tables and figures, processing handwriting, maintaining coherence across long conversations.
The limitations it lists are the ones assessment design has to work around. Probabilistic generation does not guarantee factual accuracy, and hallucination is treated as a structural property rather than a bug to be patched. Outputs are sensitive to prompt structure, so superficially similar queries yield substantially different results. Models lack genuine understanding of physical causality and embodied experience. Most consequentially, performance varies systematically with task type, complexity and domain specificity, which is what makes the cognitive gradient in the benchmark evidence meaningful rather than anecdotal.
Where AI enters the educator's own workflow
Four applications are reviewed. Prompt engineering is treated as a prerequisite skill rather than an optional refinement: the P.R.O.M.P.T. framework organizes a prompt around goal, role, output format, audience, presentation and tone, and is supplemented by retrieval-augmented generation, where uploaded course materials ground item generation in course-specific content instead of generic knowledge, and by low temperature settings (0.1 to 0.3) where consistency matters more than variation.
Automated scoring is presented with its history in automated essay scoring and its advance: LLMs can apply complex multi-dimensional rubrics without task-specific training data, and their main benefit is scalability and timeliness for formative feedback on open-ended responses from large cohorts. The stated risks match those documented elsewhere in this wiki — opacity, inexactness, and the legal questions raised by transmitting student work to commercial platforms. Formative Feedback is configured against established criteria from Hattie and Timperley (2007): criterion-referenced, precise, focused on what matters, delivered supportively.
Item construction is where the paper is most concrete. Multiple-choice items remain central to large-cohort examinations, and constructing good ones is demanding technical work: specified objectives, unambiguous stems, exactly one defensible answer, plausible distractors drawn from typical Misconceptions about AI, and calibration to a cognitive level. An agentic workflow of five stages — objective analysis, item construction, distractor development, quality self-review, difficulty estimation — is offered as a route to that work at scale, with the explicit condition that subject-matter experts review every generated item before operational use.
Recommendations for exam design
The recommendations follow from the cognitive gradient. Assessment should be shifted from product-oriented to process-oriented formats that document the writing and thinking behind a submission, through successive drafts, reflective journals on AI use and critical annotations of AI-generated content. It should require personalized, situated and experientially grounded responses — applying a theoretical framework to the student's own professional experience, or analyzing data collected in their own institutional context — because these embed a personal knowledge component an LLM cannot supply. Oral and interactive formats retain a validity advantage precisely because they require real-time performance that an examiner can probe, though the paper notes that real-time LLMs delivered through wearable devices are beginning to erode even that advantage. Practical and experimental tasks requiring physical manipulation, observation or hands-on construction remain beyond text-based systems.
Alongside the redesign sits disclosure and transparency: AI is to be used openly in the educator's workflow and documented in the student's, and the paper's framing of the widened competence set is AI literacy rather than tool avoidance — students moving from the role of author to the role of editor who evaluates accuracy, identifies logical gaps, adapts texts to disciplinary conventions and exercises judgment about what constitutes a good argument.
Integrity, validity and legal risks
The paper separates academic integrity from detection technology and argues that enhanced detection and penalty is not the appropriate response. The more fundamental question is what authentic intellectual contribution looks like in an AI-saturated environment, and how assessment can credibly evidence it, which it calls an educational question requiring educational answers. Its position is consistent with the institutional drift it describes: the market for online proctoring is projected to grow, yet guidance at many universities is moving toward assessment redesign rather than AI detection, and online proctoring is itself criticized on grounds of data protection, proportionality, algorithmic bias and student rights.
Legally, it flags that AI use for high-stakes grading requires scrutiny under data protection rules and copyright considerations when student work is transmitted to commercial platforms, and that the EU Artificial Intelligence Act, in force since 2024 and phased in over several years, designates certain educational AI applications as high-risk systems subject to enhanced transparency and accountability. Institutionally, most higher education institutions currently permit AI in assessment only in an advisory or supplementary capacity, with the final grading decision remaining with a qualified human.
The paper's account of its own limitations
The paper is a conceptual review and says so: no new data were collected, and the empirical evidence discussed is cited from other work. Its own stated limitations are that the field moves fast enough to make any systematic account of LLM capabilities potentially outdated within months, since the cited findings reflect specific model versions at specific times; that the practical guidance is illustrative rather than prescriptive and will vary by discipline, institution and pedagogical context; and that evidence on the long-term consequences of AI tool use — for foundational writing skills, critical thinking and academic self-efficacy — remains limited. It also carries a clear provenance caveat: it is a self-described preprint that has not been peer reviewed, and the author discloses that a large language model was used for manuscript preparation including text restructuring, language editing and reference formatting, while the intellectual content and substantive claims are the author's own. Readers should treat the framework as a synthesis of existing evidence with a design argument attached, not as validated by new findings.
What this means for practice
- Instructors. Re-target examination items at the Analyze level and above, where documented model performance falls — GPT-4 0.76 at Analyze against 0.99 at Remember, Llama 3 at 0.59 — and require students to document the thinking behind a submission through successive drafts, reflective journals on their AI use, or critical annotations of AI-generated content.
- Assessment designers. Build the assignment in one of the four AI-robust format families: practical and experimental tasks, oral and interactive formats such as viva voce with examiner questioning, contextually situated tasks anchored in the student's own placement or data, and process-oriented portfolio and draft sequences.
- Educators. Use the same models in your own workflow with the conditions the paper sets: prompt engineering organized around goal, role, output format, audience, presentation and tone, item generation grounded with retrieval-augmented generation on your own course materials, low temperature settings (0.1–0.3) where consistency matters, and subject-matter expert review of every generated item before operational use.
- Administrators. Keep the final summative grading decision with a qualified human examiner, and treat transmission of student work to commercial scoring platforms as a data-protection and copyright decision rather than a procurement detail, since the paper notes most institutions currently permit AI only in an advisory or supplementary capacity.
Limitations
- The paper is a conceptual review: no new data were collected, and every empirical claim is cited from other work rather than re-tested in it, so the framework is a synthesis of existing evidence with a design argument attached.
- It is a self-described preprint that has not been peer reviewed, and the author discloses that a large language model (Claude) was used for manuscript preparation including text restructuring, language editing and reference formatting, with the substantive claims retained as the author's own.
- The capability account rests on specific model versions at specific points in time — the Huber and Niklaus mapping of 43 Benchmark tasks, for example — which the author states can be outdated within months of publication.
- Evidence on the long-term consequences of AI tool use for foundational writing skills, critical thinking and academic self-efficacy remains limited, and the practical guidance is illustrative rather than prescriptive across disciplines, institutions and pedagogical contexts.
Connected Concepts
- Assessment — the object of the redesign argument and the paper's central term
- Authentic Assessment — the AI-robust direction of travel: contexts, processes and performances AI cannot supply
- Assessment Validity — what is at stake when an AI can complete a task indistinguishably from a student
- Academic Integrity — reframed away from detection and penalty toward what authentic contribution looks like
- Automated Assessment — AI-assisted scoring, its scalability benefit and its black-box risk
- Summative Assessment — high-stakes grading where the paper keeps the human examiner responsible
- Formative Assessment — the case where timeliness of AI feedback is most clearly a pedagogical gain
- Feedback — configured against criterion-referenced feedback principles
- Large Language Models (LLMs) — the technology whose documented limitations structure the whole design argument
- Generative AI — treated as a condition of assessment rather than an intruder into it
- Prompt Engineering — P.R.O.M.P.T., retrieval-augmented generation and low-temperature settings as prerequisites
- Remote Proctoring — criticized on data protection, proportionality, bias and student-rights grounds
Connected Articles
- Beyond Detection: Redesigning Authentic Assessment in an AI-Mediated World — The move beyond detection toward authentic assessment design
- Reconsidering the Use of Oral Exams and Assessments: An Old Way to Move Into a New Future — Oral examinations as an AI-robust assessment format in practice
- From Cognitive Outsourcing to Reallocation: A 3P Analysis of Student–Generative AI Engagement in Unsupervised Assessments — Cognitive outsourcing and what GenAI does to assessment evidence
- LLMs Do Not Grade Essays Like Humans — The rater-bias finding the paper cites as an automated scoring limitation
- Assessing the Quality of AI-Generated Exams: A Large-Scale Field Study — AI-generated exam items evaluated for quality in a real examination setting
- Ensuring Academic Integrity through Automated Online Exam Proctoring: A Decade-Long Systematic Review — A decade of proctoring evidence, the technology the paper treats as contested
- Can AI Evaluate Assessment? A Study of Large Language Model Meta-Assessment Performance — Whether LLMs can judge the quality of assessments themselves
- A Human-in-the-Loop Framework for AI-Assisted Scoring in Large-Scale Writing Assessment — Keeping a human examiner in the grading loop at national scale
- Asynchronous Oral Assessments: Enhancing Integrity, Engagement, and Communication in the AI Era — Scaling the oral formats the paper recommends beyond live viva voce
- Responsible Assessment in the AI Era: Key Insights from a Future-Focused Conference — Institutional framing for responsible assessment in an AI era
Citation
Klapproth, F. (2026). AI-assisted assessment and instruction in higher education: Foundations, applications, and implications for exam design. PsyArXiv Preprints (preprint, not peer reviewed).