On this page

Process-oriented assessment — the design of Assessment around how a learner reaches a conclusion rather than only around the artifact they hand in: staged submissions, reflective justification, dialogic and oral exchange, draft and version history, and criteria students can apply to themselves. It is not a synonym for task realism and not a claim about who marks the work: a realistic task whose only evidence is one unsupervised product is exposed to the same substitution problem as the essay it replaced. Generative AI is what makes the distinction urgent, because fluent output can now be produced on demand and a finished file no longer carries information about whose reasoning produced it.

Questions to Consider

  • If a student submits a polished essay, what does that artifact actually let you claim about their thinking — and which claims does it only appear to support?
  • Process evidence asks for drafts, justification, and dialogue. Who in your setting is fluent in that kind of academic self-articulation, and who might be judged on their comfort with it rather than on their understanding?
  • Detection is unreliable and redesign around process is expensive. Where would you spend a limited institutional budget, and what evidence would change your answer?

Introduction

Assessment in higher education has long inferred the thinking from the product. A submitted essay, report, or examination answer stood in for the reasoning behind it, and grading treated a well-argued artifact as evidence of a well-argued process. Thapa and Lewis (2026) name the assumption that generative AI breaks: a polished response "may no longer represent anyone's thinking process," so the inference from artifact to understanding stops being safe. Their answer is not a better detector but a different evidence base. Process-oriented assessment collects evidence of interpretation, justification, and knowledge construction as they develop, so that reasoning stays visible when output can be produced on demand.

It helps to say what the concept is not. It is not authentic assessment, which asks whether a task mirrors worthwhile real-world practice; Thapa and Lewis keep the axes separate and insist that an authentic task with a single unsupervised product is still exposed. It is not AI detection, which asks whether the tool was used rather than how a decision was made. And it is not automated assessment, which concerns who or what does the marking: a machine-graded staged portfolio can be process-oriented, and a human-graded final essay need not be. The concept concerns which evidence is collected, over what period, and under what conditions it is produced.

The idea has moved from a proposition to a recurring recommendation across the knowledge base. It appears as a named priority in AIED's agenda for learner agency and Motivation, as the organizing research agenda for the retention and transfer problem in performance-versus-learning research, as "Black Box Assessment" in Brunnström and Palmqvist's account of self-regulated learning with AI, and as a request from practicing teacher educators who found their take-home written work could no longer evidence learning at all.

What process evidence looks like in practice

The most explicit design vocabulary comes from Thapa and Lewis, whose four commitments carry the argument into task design: reflective justification, staged task design, dialogic engagement, and evaluative transparency. In practice, students explain interpretive decisions rather than only presenting conclusions, submit work that develops across checkpoints, take part in oral or dialogic exchanges about their reasoning, and work against criteria they can apply themselves — the terrain of feedback literacy and Self-Assessment. Boysen's (2026) proposal for open learning practices translates the same logic into an audit trail: a shared learning plan set before the project begins, a learning analysis plan naming required and forbidden sources, and documented drafts and version histories, preserved through Track Changes and saved file versions, that expose the work process and explain what informed each revision. He frames this as self-regulation made visible.

Practitioner and framework accounts fill in the artifacts. Uden and Hwang's (2026) LEARN framework, aimed at problem-based learning Assessment, asks for process-focused rubrics, oral justifications, staged submissions, and reasoning traces with AI-use disclosure, and treats reflection through learning journals and think-aloud protocols as the metacognitive engine of the sequence. Moganadas and colleagues (2026) make AI-transparent, process-oriented assessment one of five researchable propositions, requiring documentation, verification, and reflective justification rather than prohibition or detector use. Padhy (2026) recommends capturing revision history, interaction patterns, and time-on-task, paired with transparent disclosure. Goldstein, Marae-Haj and Zidan (2026) report the same artifacts emerging from practice after a Trust crisis: prompting records, document version history, monitored group contributions, and reflective "journey journals" kept in physical notebooks — alongside the plain finding that take-home written work could no longer evidence learning.

A different axis from authenticity, and from marking

Holding the axes apart matters, because conflating them is how a redesign looks done while the exposure remains. Sharma (2026) makes the complementary distinction between authenticity and integrity-oriented design: authenticity asks whether a task mirrors worthwhile practice, while integrity-oriented design asks whether learners can justify their decisions and assume responsibility in relation to disciplinary standards, which is why the practices he offers — annotated decision trails, verification of GenAI-contributed claims, oral defense, draft differences with version history — are selected for the judgment they make visible rather than for their realism. Espino and Espino's (2026) ten-year review of business education reaches the same design conclusion from the literature side, finding the field converging on authentic, process-visible tasks that require interpretation, contextual application, and iterative refinement, and treating isolated tool adoption inside unchanged assessment structures as its most persistent failure.

The construct is also distinguishable from the question of how evidence is scored. Validity reasoning is where the two meet: Brunnström and Palmqvist (2026) argue that when a chatbot can produce a plausible examination answer, the submitted product becomes a weaker indicator of what the student learned — an assessment-validity problem rather than only an integrity problem. That is the same inference at risk in Yan, Greiff, Lodge and Gašević's performance-versus-learning argument, where performance gains that shrink once assistance is withdrawn show a gain living in the tool rather than in the learner, and where the proposed remedy is process measures such as retention and transfer tests rather than immediate task success.

Why detection cannot carry this load

Process-oriented assessment is usually proposed as the alternative to detection, and the evidence about detection explains why. The detection literature finds error rates, task-dependence, and patterned bias against non-native and disabled writers, and the most recent work lands on a procedural conclusion: a detector score can prompt closer review but is not a finding. Munoz and colleagues (2026) show what that costs in practice. Coding 1,162 generative-AI misconduct allegations, they found the strongest evidence types are generated by the investigation or by supervision, that process evidence such as drafts, supervision meetings, and presentations appears only where those practices already exist — so in an unredesigned assessment it is simply absent — and that detector output drew the lowest probative ratings of any category. The design consequence is ordering: process evidence has to exist before a case needs it, and Weidlich's (2026) reasoning about which inference is at risk makes the same point from the validity side, since attempts to restore assessment security through detection introduce construct-irrelevant variance rather than evidence about the learner.

The design and workload costs

The sources are unusually candid that this is expensive. Thapa and Lewis state the equity caution directly: reflective and dialogic tasks can privilege students who are more confident in academic self-articulation unless they are inclusively designed and scaffolded, and the relational, feedback-intensive work intensifies the emotional and professional labor of educators in large or resource-constrained settings — a contradiction with institutional expectations for scalability, standardization, and measurable performance. Sharma adds the interpretive version: requiring documented reasoning privileges Learners more fluent in reflective discourse, and judgment as evidence "remains relational and situated rather than mechanically verifiable."

The operational costs are itemized elsewhere. Boysen lists new technology, revised assignments and syllabi, teaching students unfamiliar skills, and extra material to collect and grade, warns that low-stakes credit for process can inflate grades, and concedes the model suits large skill-development projects far better than day-to-day knowledge-acquisition work; he also notes that process tracking is what some critics call AI surveillance. Goldstein, Marae-Haj and Zidan's participants reported that establishing process documentation took real work, that nine of thirteen funded advanced tools from their own pockets in a way that could deepen the digital divide, and that several fell back on supervised examinations and anticipated oral thesis defenses. The equity worry has an AI-specific edge in Brunnström and Palmqvist's finding that unguided tool use demands an interaction-management skill that is unevenly distributed, so AI "may be most beneficial to already advantaged students" — the same distributional risk that attends asking students to manage a documented process. Lopez-López, Bru-Cordero and Correa-Álvarez (2026) found that students who invested more independent study time judged AI use more strictly, and recommend testing whether disclosure templates, oral defenses, and process-based assessment reduce ethical uncertainty — a study the field still needs.

What remains unresolved

Thapa and Lewis's account is conceptual: no participants, no data collection, and no implementation trial, so the four principles are argued rather than tested, and the paper itself calls for empirical work on student experience, educator workload, and long-term viability. Feasibility is left undifferentiated across settings — transfer to large lectures, laboratories, studios, or clinical placements is not addressed. Lane's (2026) position paper notes the institutional ceiling: process-based assessment is blunted by grading and transcript structures that reward polished final artifacts, so classroom redesign alone does not change the incentive. Measurement is thin as well: Jin and colleagues' (2026) review of AI-literacy instruments calls for performance-based and process-oriented assessment and behavioral-trace measurement of prompting and verification, but reports that the existing instruments lean on self-report. No source here reports a controlled trial of a process-oriented redesign against a product-oriented one, which leaves the central claim — that process evidence is both harder to counterfeit and fairer to judge — argued from principle, practitioner report, and case-file evidence rather than demonstrated at scale.

Connected Concepts

Connected Articles

Connected Resources

  • Process Feedback
    A free, process-based alternative to AI detection: it records how a piece of writing was produced and turns that into a report for a conversation about learning rather than a verdict.

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.