Concept
Assessment
Assessment — the process of gathering and interpreting evidence about what learners know and can do, and the methods used to evaluate learning. AI in education has fundamentally reshaped assessment: it powers automated grading and scoring, generates and adapts assessment items, and raises deep questions about what assessments actually measure when students can use AI. Assessment is the umbrella concept that organizes the knowledge base's coverage of formative assessment, automated grading, validity, and educational measurement.
Questions to Consider
- Assessment here is defined as gathering and interpreting evidence about what learners know and can do. Before reading, how would you finish the sentence 'an assessment is valid if…' — and what would change about that answer if every student could secretly use AI to produce their work?
- The page claims that how a student uses GenAI (evaluative integration versus uncritical shortcut uptake) predicted performance, while how often they used it predicted nothing. What does this suggest about a common instinct to regulate AI by limiting its use?
- A learner may deliver professional-standard work they cannot reproduce without the tool — severing the inference between performance and underlying ability. Have you encountered a 'competency' being awarded for work that the student couldn't actually do unaided? What did that reveal?
- The constructive question the page offers is not 'how do we stop students using AI?' but 'how do we enable thoughtful use in contexts that mirror their future work?' What would assessment look like in your field if that were the goal?
- The DRIVE framework suggests assessing the quality of a student's engagement with GenAI — not just the artifact — by looking at whether they steer prompts strategically and integrate their own ideas. What would you look at to tell a deep, reflective AI interaction from surface consumption?
- AI-mediated assessment is diversifying into oral exams, portfolios, and conversational formats that reduce anxiety and feel professionally relevant. Which assessment format from your own experience do you think is most 'AI-resistant' — and is resistance the same thing as educational value?
Introduction
Assessment is central to AI in education for two reasons. First, AI itself is used to assess students — grading essays, code, short answers, and exams at scale. Second, AI in the classroom changes what assessments can validly measure, since students may use generative AI to produce work. The field therefore spans both the tools that automate assessment and the validity and integrity questions that AI raises.
How AI is used in assessment
- Automated assessment: AI-based assessment spans multiple modalities — multiple-choice, short answer, essay, code, and performance-based evaluation — through automated grading, automated essay scoring, and automated question generation.
- Formative assessment: AI systems generate, validate, and adapt formative assessment items at scale, informing ongoing instruction rather than only summative evaluation.
- Learning analytics and measurement: Learning Analytics and Educational Measurement connect assessment data to learning processes, using Item Response Theory, Knowledge Tracing, and Student Modeling to interpret performance. Razavi and Powers (2026) demonstrate LLM-based item-difficulty estimation as a measurement input: across 5,170 K-5 math and reading items calibrated under the Rasch IRT model, GPT-4o's zero-shot ratings correlated moderately-to-strongly with true difficulties (r = 0.83 math, r = 0.81 reading) but varied by grade, while a feature-based approach (LLM-extracted features into tree-based models) reached correlations up to r = 0.87. The study offers a practical seven-step workflow for testing professionals, while cautioning that generalizability beyond K-5 math and reading is unclear.
- Feedback loops: AI assessment increasingly feeds into feedback loops that close the cycle from assessment to learning.
- E-portfolio assessment: e-portfolios assemble student work and reflections over time as a process-based, AI-robust assessment form. Generative AI can assist the portfolio process — generating feedback, Scaffolding reflection, and (with appropriate rubric design) supporting evaluation — while the portfolio's emphasis on reasoning traces and drafts resists AI fabrication. AI-assisted portfolio assessment and ChatGPT + e-portfolio for EFL speaking show AI can enhance both the portfolio experience and learner feedback literacy.
- Open-ended grading reliability is model-dependent: Pecuchova, Benko & Drlik (2025) benchmarked eleven GenAI and sentence-embedding models against two expert graders on 1,885 open-ended software-engineering responses: only GPTo1 reached almost-perfect agreement (Fleiss' Kappa 0.82), while reference-based models penalized correct-but-divergently-phrased answers. Because GPTo1 was the only model deemed deployable without oversight but carries proprietary API costs, the authors recommend hybrid strategies pairing advanced models with affordable options or human oversight in resource-constrained settings. A PRISMA-guided systematic review of 42 empirical studies (2023–2025) corroborates this conditional-reliability picture across the whole grading field: LLMs match human raters on closed-ended and short-answer tasks but cannot fully replace human judgment on complex, open-ended, or subjective work, and no uniform grading bias emerged — models were sometimes more lenient, sometimes stricter, and often avoided extreme scores (Jukiewicz Chatgpt Teacher Assessment Feedback 2026). Iterative rubric refinement can push LLM open-ended scoring toward near-human reliability in high-stakes medical settings: Olvet et al. (2026) found that once faculty repeatedly revised analytic and holistic rubrics on error-pattern analysis, GPT-4 reached substantial-to-almost-perfect agreement with faculty graders on three of four pre-clerkship questions (weighted kappa up to 0.94), while the residual holistic-rubric item stayed at only moderate agreement (κw = 0.54) — showing both the payoff of human-in-the-loop rubric engineering and its limits on synthetic, holistic tasks.
Validity and measurement challenges
AI raises fundamental validity questions: do AI-graded assessments measure student learning or AI-prompting skill? Does student use of AI invalidate traditional assessments? Key challenges include:
- Construct validity: Research on competency-based education shows generative AI has severed the inference between performance and underlying ability — a learner may deliver professional-standard work they cannot reproduce without the tool. This motivates reconceptualising what competencies are assessed.
- Coauthorship and integrity: Coauthorship integrity proposes a new source of validity evidence violated when students submit AI-generated content they do not understand, and explores conversational "AI Vivas" as a response.
- Psychometric quality: psychometrically aware AI and confidence-aware assessment work to keep AI scoring reliable, unbiased, and interpretable.
- Evaluation of AI assessors: AI ed evaluation provides the methods and benchmarks for determining whether automated assessors actually work — on reliability, Pedagogy, and equity — rather than headline accuracy alone.
- Quality-dependent reliability of AI and peer grading (2026): Comparing ChatGPT, peer, and instructor grading of the same undergraduate group projects, Usher & Faraon (2026) showed grading alignment with the instructor is conditional on the quality of student work: ChatGPT inflated low-quality submissions most and aligned better with the instructor on high-quality work, while peers aligned best on weaker work and under-graded strong projects. The finding challenges the binary "reliable vs unreliable" framing, pointing toward conditional-reliability models in which alternative assessors are matched to task and performance level.
- Human grading is itself value-laden (2026): Luo & Dawson (2026) show that even human grading of GenAI-assisted work is not a neutral, criteria-based act. Scenario-based interviews with 33 university teachers revealed grading decisions driven by person-oriented (honesty, diligence), capability-oriented (independence, GenAI skill, disciplinary mastery), relation-oriented (trust), and justice-oriented (fairness, beneficence) values — extending beyond the assignment to teachers' conjecture about the student. The study reframes the assessment question from "is GenAI use cheating?" to "how do teachers' value judgements shape grades, and are those values relevant to the outcomes being assessed?" — foregrounding validity and "two-way transparency" about how GenAI use will affect grades.
Integrity and the debate over detection
AI in assessment has intensified the integrity conversation. One strand focuses on detecting AI-generated text, while a growing body of research argues that detection is a limited, situational tool — not a strategy of first resort. Beyond Detection and Responsible Assessment argue that authenticity cannot be policed into existence; it must be redesigned, positioning AI as a declared collaborator and prioritizing authentic, process-based assessment over surveillance. Walton et al. (2025) ground this in evidence of how students actually judge their way through assessment with GenAI: scroll-back interviews with 26 students revealed a spectrum of six judgement events — from critically evaluating AI knowledge and learning through AI's limitations, to adopting ideas uncritically and misjudging AI contributions as their own. Stamatoulis et al. (2026) add a quantitative counterpart: across 157 students, how GenAI is used (evaluative integration to support understanding vs. low-verification shortcut uptake) predicted performance, while simple usage frequency predicted neither performance nor academic Self Efficacy. Together these studies reframe the assessment question from whether students use AI to how they judge and pattern that use.
Assessment redesign in the AI era
The constructive question in the knowledge base's assessment literature is not "how do we prevent students from using AI?" but "how do we enable them to use it thoughtfully in contexts that mirror their future work?" This reframes assessment around:
- Authentic and process-based tasks that make AI use visible and assessed. Authentic assessment — examining student performance on worthy, realistic tasks — is the leading response to AI's challenge: any task an LLM can credibly simulate loses its validity, so authenticity must be redesigned around real-time collaboration, digital and social contribution, and individual meaning-making. This connects to design frameworks for authentic assessment, authentic products and authenticated processes, and tool-invariant assessment of process.
- Responsible assessment design grounded in validity evidence (Responsible Assessment AI Era Stanford 2026)
- Task-level AI permissions derived from assessment targets: McCorkle (2025) shows the assessment-design work that precedes an AI policy — inventorying every task in a project, specifying what is being assessed and against which objective, and permitting or prohibiting AI per task on that basis (brainstorming and image curation allowed; composing learning objectives and slide design not). The same alignment exercise doubles as a check on the inference the assessment supports, because it forces the instructor to name the performance that a grade is meant to warrant (Assessment Validity).
- Coauthorship and declaration as part of the assessment contract
- Production as a competency — evaluating learners' ability to direct tools and produce professional-standard work (Competency Based Education GenAI Production 2026)
- Assessing the interaction process, not just the artifact — the DRIVE framework (Directive Reasoning Interaction + Visible Expertise) treats the quality of a student's engagement with GenAI as the assessed construct. It distinguishes surface consumption from deep, reflective interaction by looking at whether students steer prompts strategically (DRI) and integrate and develop their own disciplinary ideas through the exchange (VE), grounding process-focused criteria in theories of self-directed learning and cognitive engagement along the lines of the ICAP hierarchy. This makes DRIVE an example of AI-mediated authentic assessment — a rubric for evaluating how learners partner with GenAI rather than a detection tool.
A proposal in this literature pushes past redesign-within-the-current-frame. El Khoury and Ma (2026) argue that reform organized around preventing misconduct or detecting AI use narrows the educational imagination to control and compliance, and propose joyful assessment instead: assessment that is safe, emotionally responsive, empowering and supportive of student student agency, with safety as the load-bearing condition because without it emotional attunement becomes performance, empowerment becomes pressure and agency becomes risk. Their framing inverts the detection agenda — integrity becomes a consequence of designing assessment students want to engage in rather than its starting point — and they position instructor-built AI agents (custom GPTs, Gems, Copilot Studio agents) as a low-stakes rehearsal space where students practise before judgement, with the claim that the AI organizes evidence while the instructor interprets it.
Implications for AI in education
-
Assessment and learning are inseparable: good AI assessment should support learning (formative feedback) as much as it evaluates it.
-
Validity must be reconceptualised: when AI can produce student work, assessments must measure processes, judgment, and authentic production, not just outputs.
-
Automation must be evaluated rigorously: automated assessors need psychometric and fairness evaluation, not just accuracy claims.
-
Integrity shifts from detection to design: the most robust response to AI in assessment is designing tasks where AI use is expected, declared, and scrutinised.
-
Innovative practices can address multiple problems at once. Mesny, Roberge-Maltais & Galy (2026) argue that a set of five mutually reinforcing practices — authentic assessment, self- and peer-assessment, reassessment, standards-based grading, and ungrading — aligned with the "assessment for learning" paradigm can counter the harms of traditional, summative-heavy, norm-referenced grading (superficial learning, eroded intrinsic motivation, stress and anxiety, inequity, and compromised integrity) in the generative AI era. They find uptake is uneven across fields — self- and peer-assessment dominate while the grading-focused innovations remain marginal — and urge educators to engage more actively and reciprocally with assessment and grading innovation, backed by incremental experimentation and institutional support.
-
AI-mediated assessment is diversifying. AIvaluate shows an LLM-augmented conversational agent reduced student anxiety during performance-based assessments; Pentland (2026) finds asynchronous oral assessments offered higher engagement and were perceived as professionally relevant; graph-based ITS uses adaptive knowledge-state tracking to inform assessment.
Connected Concepts
- Pedagogical Partnerships — Pedagogical Partnerships
- Formative Assessment — Formative assessment: AI-generated, validated, adaptive items at scale
- Automated Assessment — Automated grading and scoring across assessment modalities
- Authentic Assessment — Process-based, AI-robust authentic tasks
- Assessment Validity — Validity of assessments under generative AI
- Educational Measurement — Measurement theory underpinning AI assessment
- Automated Essay Scoring — Automated essay scoring
- Automated Question Generation — Automated question generation
- Summative Assessment — Summative assessment: AI-resistant formats (oral, proctored, closed-book exams)
- Item Response Theory — IRT for interpreting AI-era assessment responses
- Psychometrically Aware AI — Psychometrically aware AI scoring
- Learning Analytics — Analytics connecting assessment data to learning
- Feedback — Feedback loops closing the assessment-to-learning cycle
- Feedback Literacy — Learner feedback literacy
- Academic Integrity — Integrity and the debate over AI detection
- AI Detection — Detecting AI-generated text
- AI Ed Evaluation — Methods and benchmarks for evaluating automated assessors
- Eportfolio — Process-based e-portfolio assessment
Connected Articles
-
AI Agents Joyful Assessment Third Space 2026 — AI agents, joyful assessment, and third space
-
Mccorkle Aligned GenAI Course Policy 2025 — Task-level AI permissions derived from what is assessed (McCorkle 2025)
-
LLM Comparative Judgment Writing Screening 2026 — Validity of Large Language Model Comparative Judgment for Universal Writing Screening
-
Student Attention Estimation Fairness 2026 — Fairness-Aware Multimodal Transformer Modeling for Real-Time Student Attention Estimation
-
Usher Faraon Who Grades Best 2026 — Comparing ChatGPT, peer, and instructor grading across project quality levels (Usher & Faraon 2026)
-
Biology Grade Vulnerability GenAI 2026 — Vulnerability of biology course grades to AI-mediated dishonesty (Chan et al. 2026)
-
Evaluation Age AI Output Evidence 2026 — Evaluation in the Age of AI
-
Causal Modelling Competency Assessment 2026 — Causal Modelling of Support Interventions for Student Competency Assessment
-
Responsible Assessment AI Era Stanford 2026 — Responsible assessment in the AI era
-
Beyond Detection Authentic Assessment AI 2025 — Beyond detection: authentic assessment
-
Coauthorship Integrity Reconceptualising Assessment Validity For The Age Of Gene — Coauthorship integrity and assessment validity
-
Competency Based Education GenAI Production 2026 — Competency-based education after generative AI
-
GenAI Assessment Governance — Evidence-centered governance of generative AI in assessment
-
LLM Difficulty Calibration Programming Exams 2026 — LLM-based difficulty calibration for programming exams
-
AI Literacy Assessment Misalignment — AI literacy and assessment misalignment
-
Hybrid E Assessment Semi Automated Grading — Hybrid e-assessment and semi-automated grading
-
Cotal Formative Assessment Scoring 2026 — Formative assessment scoring
-
Cong Confidence ASAG 2026 — Automatic short-answer grading
-
AI Assessment Scale Reform — AI assessment scale reform
-
Ithaka Sr AI Skills College Graduates 2026 — Lack of shared AI-skills assessment frameworks in higher education
-
Ssaho AI Academic Integrity Review 2025 — AI integrity review: detection must pair with assessment redesign
-
Young People Learning Generative AI Rapid Review 2026 — Evaluate learning beyond immediate GenAI-supported performance
-
Generative AI Reduced Study Time Math — Proctored, unassisted measures essential; non-proctored inflated by AI
-
Fenton Oral Exams AI Authentic Assessment 2025 — Reconsidering oral exams as authentic, AI-resistant assessment
-
Shap LLM Rationales Teaching Quality Assessment — SHAP and LLM rationales for rubric-based assessment
-
End Of Assessment AI Disruption Transformation 2026 — End of assessment: AI disruption and transformation of assessment
-
Can AI Evaluate Assessment LLM Meta Assessment 2026 — Can AI evaluate assessment? LLM meta-assessment
-
Bassett AI Detectors Education 2026 — Heads we win, tails you lose: AI detectors in education (Bassett et al. 2026)
-
Aivaluate Anxiety Assessment 2026 — AIvaluate: LLM-Augmented Assessment of Student Anxiety (2026)
-
Graph ITS Adaptive Algorithms 2026 — Graph-Based Intelligent Tutoring for Dynamic Domains (2026)
-
Asynchronous Oral Assessment 2026 — Asynchronous Oral Assessments in the AI Era (Pentland 2026)
-
Harmogen AI Assessment Rubric Generation — HARMOGEN-R: AI assessment rubric generation
-
IRT Human GenAI Mcq Responses — Using IRT to separate human and GenAI MCQ responses
-
Assessing Student Drive Framework 2025 — DRIVE: assessing learning through GenAI interaction (DRI + Visible Expertise)
-
AI Writes Code Student Writes Model 2026 — Model authorship: theory & measurement for learning-by-construction with GenAI
-
Code To Learn GenAI Artifact Construction 2026 — CtL-GenAI: constructionism framework for artifact construction
-
Dollinger Equitable Assessment AI 2026 — Equitable assessment design with AI
-
Nicola Richmond Programwide Assessment GenAI 2025 — Program-wide assessment redesign for generative AI
-
AI Grading Handwritten Physics 2026 — AI grading of handwritten physics assessments (Olympiad)
-
Chatgpt Qiskit Homework Autogradable 2026 — ChatGPT solves Qiskit homework; autogradable design
-
Credentials Carry Evidence AI Agents 2026 — Credentials that carry their evidence for AI-agent work
-
Astor Computational Thinking Meta Review 2026 — Assessment as one of five dominant CT themes
-
Xiong AI Educational Measurement Review 2026 — AI reshaping assessment practice
-
Walton Bearman Assessment Judgement 2025 — Judgement in students' work with GenAI on assessment tasks (26 students, scroll-back)
-
Stamatoulis GenAI Use Patterns 2026 — Patterns of GenAI use (evaluative integration vs low-verification uptake) and outcomes
-
Luo Dawson Value Judgements Grading 2026 — Value judgements in grading GenAI-assisted work: honesty, trust, validity, and two-way transparency (Luo & Dawson 2026)
-
Razavi Powers Item Difficulty LLM 2026 — Estimating item difficulty using LLMs and tree-based ML
-
Human Capability Test Learning Outcomes AI 2026 — A human capability test for learning outcomes in the AI era (Saleh 2026)