On this page

Assessment — the process of gathering and interpreting evidence about what learners know and can do, and the methods used to evaluate learning. AI in education has fundamentally reshaped assessment: it powers automated grading and scoring, generates and adapts assessment items, and raises deep questions about what assessments actually measure when students can use AI. Assessment is the umbrella concept that organizes the knowledge base's coverage of formative assessment, automated grading, validity, and educational measurement.

Questions to Consider

  • Assessment here is defined as gathering and interpreting evidence about what learners know and can do. Before reading, how would you finish the sentence 'an assessment is valid if…' — and what would change about that answer if every student could secretly use AI to produce their work?
  • The page claims that how a student uses GenAI (evaluative integration versus uncritical shortcut uptake) predicted performance, while how often they used it predicted nothing. What does this suggest about a common instinct to regulate AI by limiting its use?
  • A learner may deliver professional-standard work they cannot reproduce without the tool — severing the inference between performance and underlying ability. Have you encountered a 'competency' being awarded for work that the student couldn't actually do unaided? What did that reveal?
  • The constructive question the page offers is not 'how do we stop students using AI?' but 'how do we enable thoughtful use in contexts that mirror their future work?' What would assessment look like in your field if that were the goal?
  • The DRIVE framework suggests assessing the quality of a student's engagement with GenAI — not just the artifact — by looking at whether they steer prompts strategically and integrate their own ideas. What would you look at to tell a deep, reflective AI interaction from surface consumption?
  • AI-mediated assessment is diversifying into oral exams, portfolios, and conversational formats that reduce anxiety and feel professionally relevant. Which assessment format from your own experience do you think is most 'AI-resistant' — and is resistance the same thing as educational value?

Introduction

Assessment is central to AI in education for two reasons. First, AI itself is used to assess students — grading essays, code, short answers, and exams at scale. Second, AI in the classroom changes what assessments can validly measure, since students may use generative AI to produce work. The field therefore spans both the tools that automate assessment and the validity and integrity questions that AI raises.

How AI is used in assessment

  • Automated assessment: AI-based assessment spans multiple modalities — multiple-choice, short answer, essay, code, and performance-based evaluation — through automated grading, automated essay scoring, and automated question generation.
  • Formative assessment: AI systems generate, validate, and adapt formative assessment items at scale, informing ongoing instruction rather than only summative evaluation.
  • Learning analytics and measurement: Learning Analytics and Educational Measurement connect assessment data to learning processes, using Item Response Theory, Knowledge Tracing, and Student Modeling to interpret performance. Razavi and Powers (2026) demonstrate LLM-based item-difficulty estimation as a measurement input: across 5,170 K-5 math and reading items calibrated under the Rasch IRT model, GPT-4o's zero-shot ratings correlated moderately-to-strongly with true difficulties (r = 0.83 math, r = 0.81 reading) but varied by grade, while a feature-based approach (LLM-extracted features into tree-based models) reached correlations up to r = 0.87. The study offers a practical seven-step workflow for testing professionals, while cautioning that generalizability beyond K-5 math and reading is unclear.
  • Feedback loops: AI assessment increasingly feeds into feedback loops that close the cycle from assessment to learning.
  • E-portfolio assessment: e-portfolios assemble student work and reflections over time as a process-based, AI-robust assessment form. Generative AI can assist the portfolio process — generating feedback, Scaffolding reflection, and (with appropriate rubric design) supporting evaluation — while the portfolio's emphasis on reasoning traces and drafts resists AI fabrication. AI-assisted portfolio assessment and ChatGPT + e-portfolio for EFL speaking show AI can enhance both the portfolio experience and learner feedback literacy.
  • Open-ended grading reliability is model-dependent: Pecuchova, Benko & Drlik (2025) benchmarked eleven GenAI and sentence-embedding models against two expert graders on 1,885 open-ended software-engineering responses: only GPTo1 reached almost-perfect agreement (Fleiss' Kappa 0.82), while reference-based models penalized correct-but-divergently-phrased answers. Because GPTo1 was the only model deemed deployable without oversight but carries proprietary API costs, the authors recommend hybrid strategies pairing advanced models with affordable options or human oversight in resource-constrained settings. A PRISMA-guided systematic review of 42 empirical studies (2023–2025) corroborates this conditional-reliability picture across the whole grading field: LLMs match human raters on closed-ended and short-answer tasks but cannot fully replace human judgment on complex, open-ended, or subjective work, and no uniform grading bias emerged — models were sometimes more lenient, sometimes stricter, and often avoided extreme scores (Jukiewicz Chatgpt Teacher Assessment Feedback 2026). Iterative rubric refinement can push LLM open-ended scoring toward near-human reliability in high-stakes medical settings: Olvet et al. (2026) found that once faculty repeatedly revised analytic and holistic rubrics on error-pattern analysis, GPT-4 reached substantial-to-almost-perfect agreement with faculty graders on three of four pre-clerkship questions (weighted kappa up to 0.94), while the residual holistic-rubric item stayed at only moderate agreement (κw = 0.54) — showing both the payoff of human-in-the-loop rubric engineering and its limits on synthetic, holistic tasks.

Validity and measurement challenges

AI raises fundamental validity questions: do AI-graded assessments measure student learning or AI-prompting skill? Does student use of AI invalidate traditional assessments? Key challenges include:

  • Construct validity: Research on competency-based education shows generative AI has severed the inference between performance and underlying ability — a learner may deliver professional-standard work they cannot reproduce without the tool. This motivates reconceptualising what competencies are assessed.
  • Coauthorship and integrity: Coauthorship integrity proposes a new source of validity evidence violated when students submit AI-generated content they do not understand, and explores conversational "AI Vivas" as a response.
  • Psychometric quality: psychometrically aware AI and confidence-aware assessment work to keep AI scoring reliable, unbiased, and interpretable.
  • Evaluation of AI assessors: AI ed evaluation provides the methods and benchmarks for determining whether automated assessors actually work — on reliability, Pedagogy, and equity — rather than headline accuracy alone.
  • Quality-dependent reliability of AI and peer grading (2026): Comparing ChatGPT, peer, and instructor grading of the same undergraduate group projects, Usher & Faraon (2026) showed grading alignment with the instructor is conditional on the quality of student work: ChatGPT inflated low-quality submissions most and aligned better with the instructor on high-quality work, while peers aligned best on weaker work and under-graded strong projects. The finding challenges the binary "reliable vs unreliable" framing, pointing toward conditional-reliability models in which alternative assessors are matched to task and performance level.
  • Human grading is itself value-laden (2026): Luo & Dawson (2026) show that even human grading of GenAI-assisted work is not a neutral, criteria-based act. Scenario-based interviews with 33 university teachers revealed grading decisions driven by person-oriented (honesty, diligence), capability-oriented (independence, GenAI skill, disciplinary mastery), relation-oriented (trust), and justice-oriented (fairness, beneficence) values — extending beyond the assignment to teachers' conjecture about the student. The study reframes the assessment question from "is GenAI use cheating?" to "how do teachers' value judgements shape grades, and are those values relevant to the outcomes being assessed?" — foregrounding validity and "two-way transparency" about how GenAI use will affect grades.

Integrity and the debate over detection

AI in assessment has intensified the integrity conversation. One strand focuses on detecting AI-generated text, while a growing body of research argues that detection is a limited, situational tool — not a strategy of first resort. Beyond Detection and Responsible Assessment argue that authenticity cannot be policed into existence; it must be redesigned, positioning AI as a declared collaborator and prioritizing authentic, process-based assessment over surveillance. Walton et al. (2025) ground this in evidence of how students actually judge their way through assessment with GenAI: scroll-back interviews with 26 students revealed a spectrum of six judgement events — from critically evaluating AI knowledge and learning through AI's limitations, to adopting ideas uncritically and misjudging AI contributions as their own. Stamatoulis et al. (2026) add a quantitative counterpart: across 157 students, how GenAI is used (evaluative integration to support understanding vs. low-verification shortcut uptake) predicted performance, while simple usage frequency predicted neither performance nor academic Self Efficacy. Together these studies reframe the assessment question from whether students use AI to how they judge and pattern that use.

Assessment redesign in the AI era

The constructive question in the knowledge base's assessment literature is not "how do we prevent students from using AI?" but "how do we enable them to use it thoughtfully in contexts that mirror their future work?" This reframes assessment around:

  • Authentic and process-based tasks that make AI use visible and assessed. Authentic assessment — examining student performance on worthy, realistic tasks — is the leading response to AI's challenge: any task an LLM can credibly simulate loses its validity, so authenticity must be redesigned around real-time collaboration, digital and social contribution, and individual meaning-making. This connects to design frameworks for authentic assessment, authentic products and authenticated processes, and tool-invariant assessment of process.
  • Responsible assessment design grounded in validity evidence (Responsible Assessment AI Era Stanford 2026)
  • Task-level AI permissions derived from assessment targets: McCorkle (2025) shows the assessment-design work that precedes an AI policy — inventorying every task in a project, specifying what is being assessed and against which objective, and permitting or prohibiting AI per task on that basis (brainstorming and image curation allowed; composing learning objectives and slide design not). The same alignment exercise doubles as a check on the inference the assessment supports, because it forces the instructor to name the performance that a grade is meant to warrant (Assessment Validity).
  • Coauthorship and declaration as part of the assessment contract
  • Production as a competency — evaluating learners' ability to direct tools and produce professional-standard work (Competency Based Education GenAI Production 2026)
  • Assessing the interaction process, not just the artifact — the DRIVE framework (Directive Reasoning Interaction + Visible Expertise) treats the quality of a student's engagement with GenAI as the assessed construct. It distinguishes surface consumption from deep, reflective interaction by looking at whether students steer prompts strategically (DRI) and integrate and develop their own disciplinary ideas through the exchange (VE), grounding process-focused criteria in theories of self-directed learning and cognitive engagement along the lines of the ICAP hierarchy. This makes DRIVE an example of AI-mediated authentic assessment — a rubric for evaluating how learners partner with GenAI rather than a detection tool.

A proposal in this literature pushes past redesign-within-the-current-frame. El Khoury and Ma (2026) argue that reform organized around preventing misconduct or detecting AI use narrows the educational imagination to control and compliance, and propose joyful assessment instead: assessment that is safe, emotionally responsive, empowering and supportive of student student agency, with safety as the load-bearing condition because without it emotional attunement becomes performance, empowerment becomes pressure and agency becomes risk. Their framing inverts the detection agenda — integrity becomes a consequence of designing assessment students want to engage in rather than its starting point — and they position instructor-built AI agents (custom GPTs, Gems, Copilot Studio agents) as a low-stakes rehearsal space where students practise before judgement, with the claim that the AI organizes evidence while the instructor interprets it.

Implications for AI in education

  • Assessment and learning are inseparable: good AI assessment should support learning (formative feedback) as much as it evaluates it.

  • Validity must be reconceptualised: when AI can produce student work, assessments must measure processes, judgment, and authentic production, not just outputs.

  • Automation must be evaluated rigorously: automated assessors need psychometric and fairness evaluation, not just accuracy claims.

  • Integrity shifts from detection to design: the most robust response to AI in assessment is designing tasks where AI use is expected, declared, and scrutinised.

  • Innovative practices can address multiple problems at once. Mesny, Roberge-Maltais & Galy (2026) argue that a set of five mutually reinforcing practices — authentic assessment, self- and peer-assessment, reassessment, standards-based grading, and ungrading — aligned with the "assessment for learning" paradigm can counter the harms of traditional, summative-heavy, norm-referenced grading (superficial learning, eroded intrinsic motivation, stress and anxiety, inequity, and compromised integrity) in the generative AI era. They find uptake is uneven across fields — self- and peer-assessment dominate while the grading-focused innovations remain marginal — and urge educators to engage more actively and reciprocally with assessment and grading innovation, backed by incremental experimentation and institutional support.

  • AI-mediated assessment is diversifying. AIvaluate shows an LLM-augmented conversational agent reduced student anxiety during performance-based assessments; Pentland (2026) finds asynchronous oral assessments offered higher engagement and were perceived as professionally relevant; graph-based ITS uses adaptive knowledge-state tracking to inform assessment.

Connected Concepts

Connected Articles

Connected FAQs

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.