On this page

Synthesis: El Khoury and Ma (2026) argue that assessment reform in the generative AI era has been too defensive — organized around preventing misconduct and detecting AI use rather than asking what assessment could become. They propose joyful assessment: assessment that is safe, emotionally responsive, empowering, and supportive of student Agency, with safety as the foundation the other three rest on. Their second move is to position instructor-built AI agents (custom GPTs, Gems, Copilot Studio agents — tools designed for specific pedagogical purposes, not autonomous systems) as a third space between the formal curriculum and students' interests and experiences, where students can rehearse, explore, and practise before judgement. The paper is explicit that this is a conceptual proposal, not evidence: its worked example is offered as a design illustration, and the authors flag ethics, reliability, equity, and the thin evidence base as open problems.

Overview

The paper starts from a problem the knowledge base documents repeatedly: since ChatGPT, students can produce fluent academic work with minimal authorial effort, and most institutional responses have asked whether assessments can still support defensible claims about learning. Reviews of the field describe this as a "wicked problem" balancing validity, integrity, equity, workload and learning. El Khoury and Ma accept that framing but argue it is incomplete, because much of the work remains directed at preventing or detecting AI use rather than asking what assessment might become.

They point to a parallel literature that answers a different question. Compassionate assessment foregrounds trauma-informed, flexible, care-oriented practice; assessment for inclusion attends to equity and social justice in how diverse learners demonstrate learning; relational approaches reframe assessment as an ethical interaction between teacher and student rather than a measure of individual performance. Read alongside the GenAI debate, those traditions suggest the question: what would assessment look like if it were designed primarily to support engagement, agency and learning, with integrity as a consequence of that design rather than its starting point? The authors give that question practical urgency by noting that disengagement is one of the conditions under which academic dishonesty becomes more likely, while reform agendas focused on stress-testing assessments or constraining AI use leave engagement untouched.

Worth noting for readers in this knowledge base: the paper builds directly on existing scholarship rather than coining its framework from nothing. The four characteristics are drawn from prior work, and the pedagogy of joy comes from White et al.'s (2026) call for collaboration, belonging, and the freedom to learn without fear of failure; what is new here is the claim that AI agents can carry such a pedagogy into assessment practice.

Definition of AI agents used here. Custom AI assistants designed by instructors for specific educational purposes — Gems, custom GPTs, or Copilot Studio agents across platforms — built to support and scaffold specific learning tasks rather than to operate autonomously. That definition matters because it is much narrower than the agentic AI usually discussed in this knowledge base, and it is what makes the paper's claims about instructor authority coherent.

The Framework: Four Characteristics

Joyful assessment is defined as a rigorous yet humanizing approach that supports learning through emotional attunement, agency, and confidence, and interactivity. The authors are careful about what it is not: not entertaining assessment, not easier assessment, not assessment without standards. Their aim is the conditions under which meaningful challenge becomes possible — conditions that hold space for students' emotional, personal, creative and relational experience alongside the cognitive one.

  • Safe (and foundational). A safe assessment is one where students do not feel their grades, mental health or sense of self are at stake every time they attempt something — where early attempts are not harshly judged, peer competition is not constant, and students can ask the questions they actually need to ask. The authors make safety load-bearing: without it, emotional attunement becomes performance, empowerment becomes pressure, and agency becomes risk.
  • Emotionally responsive. Drawing on work on emotions in assessment for learning, on hope and pride in assessment contexts, and on the argument that higher education should attend to the emotional dimensions of assessment beyond anxiety alone, the paper treats affect as pedagogy rather than background noise — noting that ignoring it falls hardest on students whose prior assessment histories involved exclusion or marginalisation.
  • Empowering. Empowerment is treated as a design quality rather than motivational rhetoric: assessment that helps students feel more capable in the doing, not only in the grading. The cited example is learner-generated podcasts produced in pairs, where confidence grows from sustained supported engagement with material students helped shape.
  • Supporting student agency. Assessment connects to students' interests, experiences and identities, as in blogging assessments where students define aspects of their own content and explore topics they care about — assessment as a site of participation rather than only evaluation. Agency here is structural, not an add-on: students engage as learners with identities, not only as candidates for grading.

How AI Agents Are Claimed to Support It

The review deliberately excludes studies where AI agents primarily grade student work, on the stated grounds that automated grading sits outside the paper's scope. The authors' position is that assessment must remain relational — the instructor role cannot be substituted by a machine, but it can be extended and made more sustainable through careful design. The four clusters:

  • Safe spaces for rehearsal. A generative AI oral exam simulator lets students rehearse responses, work with prompts, and become familiar with the structure and expectations of a high-stakes assessment before encountering it for marks; an engineering-education agent customizes preparation activities and responds to individual needs ahead of summative evaluation. The unifying move is lowering the perceived stakes of attempting unfamiliar work, so students arrive at formal assessment having practised in private — the instructor's evaluative role intact, the conditions under which students arrive changed.
  • Emotional responsiveness through appraisal. Because agents sit outside the social hierarchies students navigate with peers and instructors, and remain available whenever a student is ready to try again, they can offer feedback without the perceived risk of peer pressure or instructor judgement, support iteration without the cost of looking unprepared, and prompt students to notice how learning feels alongside what it produces. The mechanism named is appraisal: when students can rehearse without an audience, set their own pace, and decide when to begin, a task that felt like a verdict can start to feel like a conversation they can shape. Repeated interaction is claimed to convert occasional feelings of Self Efficacy and agency into a settled, everyday part of students' approach to assessment.
  • Empowerment through repeated formative practice. An AI student assistant designed into the instructional workflow helps educators identify where students struggle and design more targeted preparation, Scaffolding and feedback. Nursing education's AISimBot lets students practise clinical communication through realistic scenarios before being assessed on those skills; in language education, real-time AI avatars give ESL students low-pressure pronunciation rehearsal and feedback. On feedback specifically, an agent used in a large online class supported students' generation and uptake of feedback where timely individualised feedback would otherwise be prohibitive, and a custom GPT delivering formative feedback on first-year physics lab reports was valued for structure, consistency and speed alongside human assessors.
  • Agency by making richer formats feasible. The access argument: multimodal, authentic and practice-based assessment supports deeper learning but is time-intensive to design and manage, so agents that help generate scenarios, adapt materials, build practice tasks and manage process-oriented evidence make those designs more feasible. Here AI does not reduce assessment to automation — it opens room for students to choose how they demonstrate what they know.

The third space claim

Read across those clusters, the authors argue agents are most valuable when they extend rather than replace relational practice, and that where design choices are made well the agent functions as a third space: a hybrid pedagogical zone, after Bhabha and Gutiérrez, between the formal curriculum and students' interests, experiences and ways of knowing. Their framing is emphatic that none of this happens automatically — each affordance requires deliberate design grounded in the instructor's judgement about what the assessment is for.

Worked Example: GACSS

The Generative AI Clinical Stakeholder Simulation is a custom GPT for a health communication course in a rural Appalachian context, offered as an illustrative design example rather than evidence of measured impact. Students hold three separate conversations with AI-played personas, pause for structured reflection between encounters, and receive a rubric-aligned evidence report in which every score links to quoted excerpts from their transcript.

  • Safety through rehearsal. Students practise difficult, culturally situated clinical conversations before meeting similar demands in clinical or community settings. The personas include a retired manual labourer with respiratory symptoms who mistrusts physicians and downplays his condition, and a parent managing cost barriers and time pressure at a sliding-scale clinic. The authors insist such personas be developed with local knowledge and reviewed for cultural accuracy, so the simulation does not treat communities as homogeneous or reduce them to fixed traits. Because the instructor frames it as formative rehearsal, anxiety becomes pedagogically useful.
  • Agency through sequencing and pacing. Students choose the order of the three personas, calibrating challenge themselves; each conversation continues until the student signals its end; a coaching pause offers a strategic prompt without leaving the scenario — help is available, but the student still decides how to proceed.
  • A different artefact of learning. Dialogue, not written prose, becomes the primary medium of demonstration, and the transcript becomes the assessment artefact. The authors do not claim this removes integrity concerns; they argue it shifts evidence away from a detached product toward an interactional record in which judgement, empathy, clarification and shared decision-making unfold in context.
  • Two-layer feedback. The agent produces a draft rubric-based evidence report on empathy and validation, patient education and clarity, and shared decision-making, each rating supported by exact transcript excerpts; the instructor then reviews the report, leads a debrief, and determines what the evidence means for that student. Their summary of the division of labour is the clearest statement of the paper's position: the AI organizes evidence; the instructor interprets it.

Implications and Open Problems

The paper's research agenda is more cautious than its framing, and several items are stated in terms this knowledge base can act on.

  • Multimodality is not automatically better. Noting the shift toward agents that respond through language, images, audio and embodied action, the authors cite evidence that multimodal conversational support improved learning outcomes and experience in a visually rich biology task, while text-only support increased perceived ease and engagement without the same learning benefit. The design implication: value depends on alignment among modality, task, feedback and evidence of learning, not technological novelty.
  • Speed is not validity. AI agents can supply rapid prompts, examples and transcript-based reports, but the authors say researchers should test whether that feedback is accurate, culturally responsive, accessible, interpretable and aligned with course outcomes — with hallucination, bias and overreliance treated as core threats to the trustworthiness of evidence rather than technical side issues.
  • Ethics as a condition, not a safeguard. Transparency, student consent, data protection, bias review, accessibility, human oversight and equitable access are framed as conditions of the framework itself, following UNESCO's guidance on human-centred, ethical and equitable GenAI use. An agent cannot support joyful assessment, they argue, if it makes students feel monitored, misrepresented, excluded, or unable to contest the feedback they receive.
  • Relationships and invisible labour. The strongest use of agents may be to make better human judgement possible — richer evidence, more focused debriefing, more rehearsal, more timely formative support — which requires designs including student voice, instructor interpretation and social context. The authors add that instructor workload needs measuring, because a tool that improves student experience while intensifying invisible labour may not be sustainable.
  • Differences that matter. Future work should examine how, for whom and under what conditions these benefits hold, attending to student groups, disciplines, levels of study, linguistic backgrounds, disability status and technology access.

Limitations. This is a conceptual paper, and the authors say so: it is offered as "a first conceptual step" rather than empirical evidence. The GACSS example is explicitly an illustration, not measured impact. The AI-agent literature reviewed sits mostly in the formative and relational space because studies of agents that grade work were excluded by design, so the paper says nothing about automated scoring. The authors acknowledge upfront that agents raise legitimate concerns about ethics, reliability, equity and the strength of available evidence, and note that the framework's own characteristics — safety, emotional responsiveness, empowerment, agency — are asserted rather than measured.

Connected Concepts

Connected Articles

Citation

El Khoury, E., & Ma, X. (2026). AI Agents, Joyful Assessment, and Third Space: Rethinking Assessment in the GenAI Era. Journal of Teaching and Learning, 20(4), 230–241.

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.