Concept
Assessment
Assessment — the process of gathering and interpreting evidence about what learners know and can do, and the methods used to evaluate learning. AI in Education has fundamentally reshaped assessment: it powers automated grading and scoring, generates and adapts assessment items, and raises deep questions about what assessments actually measure when students can use AI. Assessment is the umbrella concept that organizes the knowledge base's coverage of Formative Assessment, automated grading, validity, and Educational Measurement.
Questions to Consider
- Assessment here is defined as gathering and interpreting evidence about what learners know and can do. Before reading, how would you finish the sentence 'an assessment is valid if…' — and what would change about that answer if every student could secretly use AI to produce their work?
- The page claims that how a student uses GenAI (evaluative integration versus uncritical shortcut uptake) predicted performance, while how often they used it predicted nothing. What does this suggest about a common instinct to regulate AI by limiting its use?
- A learner may deliver professional-standard work they cannot reproduce without the tool — severing the inference between performance and underlying ability. Have you encountered a 'competency' being awarded for work that the student couldn't actually do unaided? What did that reveal?
- The constructive question the page offers is not 'how do we stop students using AI?' but 'how do we enable thoughtful use in contexts that mirror their future work?' What would assessment look like in your field if that were the goal?
- The DRIVE framework suggests assessing the quality of a student's engagement with GenAI — not just the artifact — by looking at whether they steer prompts strategically and integrate their own ideas. What would you look at to tell a deep, reflective AI interaction from surface consumption?
- AI-mediated assessment is diversifying into oral exams, portfolios, and conversational formats that reduce anxiety and feel professionally relevant. Which assessment format from your own experience do you think is most 'AI-resistant' — and is resistance the same thing as educational value?
Introduction
Assessment is central to AI in education for two reasons. First, AI itself is used to assess students — grading essays, code, short answers, and exams at scale. Second, AI in the classroom changes what assessments can validly measure, since students may use Generative AI to produce work. The field therefore spans both the tools that automate assessment and the validity and integrity questions that AI raises.
How AI is used in assessment
- Automated assessment: AI-based assessment spans multiple modalities — multiple-choice, short answer, essay, code, and performance-based evaluation — through automated grading, Automated Essay Scoring, and Automated Question Generation.
- Formative assessment: AI systems generate, validate, and adapt formative assessment items at scale, informing ongoing instruction rather than only summative evaluation.
- Learning analytics and measurement: Learning Analytics and Educational Measurement connect assessment data to learning processes, using Item Response Theory, Knowledge Tracing, and Learner Modeling and Adaptive Instruction to interpret performance. Razavi and Powers (2026) demonstrate LLM-based item-difficulty estimation as a measurement input: across 5,170 K-5 math and reading items calibrated under the Rasch IRT model, GPT-4o's zero-shot ratings correlated moderately-to-strongly with true difficulties (r = 0.83 math, r = 0.81 reading) but varied by grade, while a feature-based approach (LLM-extracted features into tree-based models) reached correlations up to r = 0.87. The study offers a practical seven-step workflow for testing professionals, while cautioning that generalizability beyond K-5 math and reading is unclear.
- Feedback loops: AI assessment increasingly feeds into feedback loops that close the cycle from assessment to learning.
- E-portfolio assessment: e-portfolios assemble student work and reflections over time as a process-based, AI-robust assessment form. Generative AI can assist the portfolio process — generating feedback, Scaffolding reflection, and (with appropriate rubric design) supporting evaluation — while the portfolio's emphasis on reasoning traces and drafts resists AI fabrication. AI-assisted portfolio assessment and ChatGPT + e-portfolio for EFL speaking show AI can enhance both the portfolio experience and learner Feedback Literacy.
- Open-ended grading reliability is model-dependent: Pecuchova, Benko & Drlik (2025) benchmarked eleven GenAI and sentence-embedding models against two expert graders on 1,885 open-ended software-engineering responses: only GPTo1 reached almost-perfect agreement (Fleiss' Kappa 0.82), while reference-based models penalized correct-but-divergently-phrased answers. Because GPTo1 was the only model deemed deployable without oversight but carries proprietary API costs, the authors recommend hybrid strategies pairing advanced models with affordable options or human oversight in resource-constrained settings. A PRISMA-guided systematic review of 42 empirical studies (2023–2025) corroborates this conditional-reliability picture across the whole grading field: LLMs match human raters on closed-ended and short-answer tasks but cannot fully replace human judgment on complex, open-ended, or subjective work, and no uniform grading bias emerged — models were sometimes more lenient, sometimes stricter, and often avoided extreme scores (Can ChatGPT Replace the Teacher in Assessment? A Review of Research on the Use of Large Language Models in Grading and Providing Feedback). Iterative rubric refinement can push LLM open-ended scoring toward near-human reliability in high-stakes medical settings: Olvet et al. (2026) found that once faculty repeatedly revised analytic and holistic rubrics on error-pattern analysis, GPT-4 reached substantial-to-almost-perfect agreement with faculty graders on three of four pre-clerkship questions (weighted kappa up to 0.94), while the residual holistic-rubric item stayed at only moderate agreement (κw = 0.54) — showing both the payoff of human-in-the-loop rubric engineering and its limits on synthetic, holistic tasks.
- AI-supported oral and performance assessment. A TVET design study moves assessment away from text-heavy formats toward interactive Oral Assessment supported by an LLM: across four cohorts, 21 of 33 learners rated the voice task realistic and none disagreed that speaking in real time reflected competence better than a written portfolio, while word counts for identical questions varied five- to eight-fold between cohorts and nine learners answered in two to thirteen words and were all correct. The system ran offline on a single laptop serving up to 12 simultaneous learners with Mistral 7B and faster-whisper, deleted recordings after 90 days, and left scoring judgments with assessors — an instance of Human-in-the-Loop assessment design for workplace-facing qualifications. (Designing AI-Supported Oral Assessment in TVET)
Validity and measurement challenges
AI raises fundamental validity questions: do AI-graded assessments measure student learning or AI-prompting skill? Does student use of AI invalidate traditional assessments? Key challenges include:
- Construct validity: Research on competency-based education shows generative AI has severed the inference between performance and underlying ability — a learner may deliver professional-standard work they cannot reproduce without the tool. This motivates reconceptualizing what competencies are assessed.
- Coauthorship and integrity: Coauthorship integrity proposes a new source of validity evidence violated when students submit AI-generated content they do not understand, and explores conversational "AI Vivas" as a response.
- Psychometric quality: Psychometrically Aware AI and confidence-aware assessment work to keep AI scoring reliable, unbiased, and interpretable.
- Evaluation of AI assessors: AI Ed Evaluation provides the methods and benchmarks for determining whether automated assessors actually work — on reliability, Pedagogies and Teaching Strategies, and Equity — rather than headline accuracy alone.
- Quality-dependent reliability of AI and peer grading (2026): Comparing ChatGPT, peer, and instructor grading of the same undergraduate group projects, Usher & Faraon (2026) showed grading alignment with the instructor is conditional on the quality of student work: ChatGPT inflated low-quality submissions most and aligned better with the instructor on high-quality work, while peers aligned best on weaker work and under-graded strong projects. The finding challenges the binary "reliable vs unreliable" framing, pointing toward conditional-reliability models in which alternative assessors are matched to task and performance level.
- Human grading is itself value-laden (2026): Luo & Dawson (2026) show that even human grading of GenAI-assisted work is not a neutral, criteria-based act. Scenario-based interviews with 33 university teachers revealed grading decisions driven by person-oriented (honesty, diligence), capability-oriented (independence, GenAI skill, disciplinary mastery), relation-oriented (trust), and justice-oriented (fairness, beneficence) values — extending beyond the assignment to teachers' conjecture about the student. The study reframes the assessment question from "is GenAI use cheating?" to "how do teachers' value judgments shape grades, and are those values relevant to the outcomes being assessed?" — foregrounding validity and "two-way transparency" about how GenAI use will affect grades.
Integrity and the debate over detection
AI in assessment has intensified the integrity conversation. One strand focuses on detecting AI-generated text, while a growing body of research argues that detection is a limited, situational tool — not a strategy of first resort. Beyond Detection and Responsible Assessment argue that authenticity cannot be policed into existence; it must be redesigned, positioning AI as a declared collaborator and prioritizing authentic, process-based assessment over surveillance. Walton et al. (2025) ground this in evidence of how students actually judge their way through assessment with GenAI: scroll-back interviews with 26 students revealed a spectrum of six judgment events — from critically evaluating AI knowledge and learning through AI's limitations, to adopting ideas uncritically and misjudging AI contributions as their own. Stamatoulis et al. (2026) add a quantitative counterpart: across 157 students, how GenAI is used (evaluative integration to support understanding vs. low-verification shortcut uptake) predicted performance, while simple usage frequency predicted neither performance nor academic Self-Efficacy. Together these studies reframe the assessment question from whether students use AI to how they judge and pattern that use.
Assessment redesign in the AI era
The constructive question in the knowledge base's assessment literature is not "how do we prevent students from using AI?" but "how do we enable them to use it thoughtfully in contexts that mirror their future work?" This reframes assessment around:
- Authentic and process-based tasks that make AI use visible and assessed. Authentic Assessment — examining student performance on worthy, realistic tasks — is the leading response to AI's challenge: any task an Large Language Models (LLMs) can credibly simulate loses its validity, so authenticity must be redesigned around real-time collaboration, digital and social contribution, and individual meaning-making. This connects to design frameworks for authentic assessment, authentic products and authenticated processes, and tool-invariant assessment of process.
- Responsible assessment design grounded in validity evidence (Responsible Assessment in the AI Era: Key Insights from a Future-Focused Conference)
- Task-level AI permissions derived from assessment targets: McCorkle (2025) shows the assessment-design work that precedes an AI policy — inventorying every task in a project, specifying what is being assessed and against which objective, and permitting or prohibiting AI per task on that basis (brainstorming and image curation allowed; composing learning objectives and slide design not). The same alignment exercise doubles as a check on the inference the assessment supports, because it forces the instructor to name the performance that a grade is meant to warrant (Assessment Validity).
- Coauthorship and declaration as part of the assessment contract
- Production as a competency — evaluating learners' ability to direct tools and produce professional-standard work (Knowledge, Skills, Attitudes, Production: Competency-Based Education After Generative AI)
- Assessing the interaction process, not just the artifact — the DRIVE framework (Directive Reasoning Interaction + Visible Expertise) treats the quality of a student's engagement with GenAI as the assessed construct. It distinguishes surface consumption from deep, reflective interaction by looking at whether students steer prompts strategically (DRI) and integrate and develop their own disciplinary ideas through the exchange (VE), grounding process-focused criteria in theories of Self-Directed Learning and cognitive engagement along the lines of the ICAP hierarchy. This makes DRIVE an example of AI-mediated authentic assessment — a rubric for evaluating how learners partner with GenAI rather than a detection tool.
A proposal in this literature pushes past redesign-within-the-current-frame. El Khoury and Ma (2026) argue that reform organized around preventing misconduct or detecting AI use narrows the educational imagination to control and compliance, and propose joyful assessment instead: assessment that is safe, emotionally responsive, empowering and supportive of student student agency, with safety as the load-bearing condition because without it emotional attunement becomes performance, empowerment becomes pressure and agency becomes risk. Their framing inverts the detection agenda — integrity becomes a consequence of designing assessment students want to engage in rather than its starting point — and they position instructor-built AI agents (custom GPTs, Gems, Copilot Studio agents) as a low-stakes rehearsal space where students practice before judgment, with the claim that the AI organizes evidence while the instructor interprets it.
Implications for AI in education
-
Assessment and learning are inseparable: good AI assessment should support learning (formative feedback) as much as it evaluates it.
-
Validity must be reconceptualised: when AI can produce student work, assessments must measure processes, judgment, and authentic production, not just outputs.
-
Automation must be evaluated rigorously: automated assessors need psychometric and fairness evaluation, not just accuracy claims.
-
Integrity shifts from detection to design: the most robust response to AI in assessment is designing tasks where AI use is expected, declared, and scrutinized.
-
Innovative practices can address multiple problems at once. Mesny, Roberge-Maltais & Galy (2026) argue that a set of five mutually reinforcing practices — Authentic Assessment, self- and Peer Assessment, reassessment, standards-based grading, and ungrading — aligned with the "assessment for learning" paradigm can counter the harms of traditional, summative-heavy, norm-referenced grading (superficial learning, eroded intrinsic motivation, stress and anxiety, inequity, and compromised integrity) in the generative AI era. They find uptake is uneven across fields — self- and peer-assessment dominate while the grading-focused innovations remain marginal — and urge educators to engage more actively and reciprocally with assessment and grading innovation, backed by incremental experimentation and institutional support.
-
AI-mediated assessment is diversifying. AIvaluate shows an LLM-augmented conversational agent reduced student anxiety during performance-based assessments; Pentland (2026) finds asynchronous oral assessments offered higher engagement and were perceived as professionally relevant; graph-based ITS uses adaptive knowledge-state tracking to inform assessment.
Connected Concepts
- Learners — Learners: the umbrella for the learner-side concepts
- Pedagogical Partnerships — Pedagogical Partnerships
- Formative Assessment — Formative assessment: AI-generated, validated, adaptive items at scale
- Automated Assessment — Automated grading and scoring across assessment modalities
- Authentic Assessment — Process-based, AI-robust authentic tasks
- Assessment Validity — Validity of assessments under generative AI
- Educational Measurement — Measurement theory underpinning AI assessment
- Automated Essay Scoring — Automated essay scoring
- Automated Question Generation — Automated question generation
- Oral Assessment — Oral assessment: live and recorded formats that resist AI substitution
- Summative Assessment — Summative assessment: AI-resistant formats (oral, proctored, closed-book exams)
- Item Response Theory — IRT for interpreting AI-era assessment responses
- Psychometrically Aware AI — Psychometrically aware AI scoring
- Learning Analytics — Analytics connecting assessment data to learning
- Feedback — Feedback loops closing the assessment-to-learning cycle
- Feedback Literacy — Learner feedback literacy
- Academic Integrity — Integrity and the debate over AI detection
- AI Detection — Detecting AI-generated text
- AI Ed Evaluation — Methods and benchmarks for evaluating automated assessors
- E-Portfolio — Process-based e-portfolio assessment
- Speech and Voice Technologies
- Peer Assessment
Connected Articles
- AI Agents, Joyful Assessment, and Third Space: Rethinking Assessment in the GenAI Era — AI agents, joyful assessment, and third space
- Designing an Aligned Generative AI Course Policy: An Equitable and Transparent Learner-Centered Approach — Task-level AI permissions derived from what is assessed (McCorkle 2025)
- Who grades best? Comparing ChatGPT, peer, and instructor evaluations across varying levels of student project quality — Comparing ChatGPT, peer, and instructor grading across project quality levels (Usher & Faraon 2026)
- Responsible Assessment in the AI Era: Key Insights from a Future-Focused Conference — Responsible assessment in the AI era
- Beyond Detection: Redesigning Authentic Assessment in an AI-Mediated World — Beyond detection: authentic assessment
- Coauthorship integrity: Reconceptualising assessment validity for the age of generative artificial intelligence — Coauthorship integrity and assessment validity
- Knowledge, Skills, Attitudes, Production: Competency-Based Education After Generative AI — Competency-based education after generative AI
- Generative AI as a Design Variable: An Evidence-Centered Framework for Principled Governance in STEM Assessment — Evidence-centered governance of generative AI in assessment
- Confidence Estimation in Automatic Short Answer Grading with LLMs — Automatic short-answer grading
- Reassessing Academic Integrity in the Age of AI: A Systematic Literature Review on AI and Academic Integrity — AI integrity review: detection must pair with assessment redesign
- Reconsidering the Use of Oral Exams and Assessments: An Old Way to Move Into a New Future — Reconsidering oral exams as authentic, AI-resistant assessment
- Heads We Win, Tails You Lose: AI Detectors in Education — Heads we win, tails you lose: AI detectors in education (Bassett et al. 2026)
- Exploring student anxiety and experience in performance-based assessments using AIvaluate: an LLM-augmented emotionally — AIvaluate: LLM-Augmented Assessment of Student Anxiety (2026)
- Intelligent tutoring in dynamic domains: a graph-based system for comparative analysis of adaptive algorithms — Graph-Based Intelligent Tutoring for Dynamic Domains (2026)
- Asynchronous Oral Assessments: Enhancing Integrity, Engagement, and Communication in the AI Era — Asynchronous Oral Assessments in the AI Era (Pentland 2026)
- Assessing students' DRIVE: A framework to evaluate learning through interactions with generative AI — DRIVE: assessing learning through GenAI interaction (DRI + Visible Expertise)
- A Decade of Reflection and Thematic Review on Artificial Intelligence's Impact on Educational Measurement — AI reshaping assessment practice
- How university students work on assessment tasks with generative AI: matters of judgement — Judgment in students' work with GenAI on assessment tasks (26 students, scroll-back)
- Same tool, different work: patterns of generative AI use and academic outcomes — Patterns of GenAI use (evaluative integration vs low-verification uptake) and outcomes
- Exploring value judgements in grading: will teachers mark down student work assisted by GenAI, and should they? — Value judgments in grading GenAI-assisted work: honesty, trust, validity, and two-way transparency (Luo & Dawson 2026)
- Estimating Item Difficulty Using Large Language Models and Tree-Based Machine Learning Algorithms — Estimating item difficulty using LLMs and tree-based ML
- Automated Grading of Open-Ended Questions in Higher Education Using GenAI Models
- Innovative assessment and grading practices in higher education: A critical exploration for management educators
- Can Generative Artificial Intelligence Reliably Score Open-Ended Question Assessments in Undergraduate Medical Education?
- Can ChatGPT Replace the Teacher in Assessment? A Review of Research on the Use of Large Language Models in Grading and Providing Feedback
- AI Can Do Your Homework. Now What? Report from an online workshop on computing assessment in the age of generative AI — AI Can Do Your Homework. Now What? Report from an online workshop on computing assessment in the age of generative AI
Connected FAQs
- How Do I Redesign Assessment So That a Grade Still Tells Me Something Defensible About What the Student Knows or Can Do?
- How Can I Reduce AI Cheating in My Course?
- How Do I Write a Course AI Policy and Communicate It to Students?
- How Should I Handle AI in Group and Collaborative Assignments?
- How Should We Design and Facilitate Asynchronous Online Courses When AI Can Do the Work?