Concept
Process-Oriented Assessment
Process-oriented assessment — the design of Assessment around how a learner reaches a conclusion rather than only around the artifact they hand in: staged submissions, reflective justification, dialogic and oral exchange, draft and version history, and criteria students can apply to themselves. It is not a synonym for task realism and not a claim about who marks the work: a realistic task whose only evidence is one unsupervised product is exposed to the same substitution problem as the essay it replaced. Generative AI is what makes the distinction urgent, because fluent output can now be produced on demand and a finished file no longer carries information about whose reasoning produced it.
Questions to Consider
- If a student submits a polished essay, what does that artifact actually let you claim about their thinking — and which claims does it only appear to support?
- Process evidence asks for drafts, justification, and dialogue. Who in your setting is fluent in that kind of academic self-articulation, and who might be judged on their comfort with it rather than on their understanding?
- Detection is unreliable and redesign around process is expensive. Where would you spend a limited institutional budget, and what evidence would change your answer?
Introduction
Assessment in higher education has long inferred the thinking from the product. A submitted essay, report, or examination answer stood in for the reasoning behind it, and grading treated a well-argued artifact as evidence of a well-argued process. Thapa and Lewis (2026) name the assumption that generative AI breaks: a polished response "may no longer represent anyone's thinking process," so the inference from artifact to understanding stops being safe. Their answer is not a better detector but a different evidence base. Process-oriented assessment collects evidence of interpretation, justification, and knowledge construction as they develop, so that reasoning stays visible when output can be produced on demand.
It helps to say what the concept is not. It is not authentic assessment, which asks whether a task mirrors worthwhile real-world practice; Thapa and Lewis keep the axes separate and insist that an authentic task with a single unsupervised product is still exposed. It is not AI detection, which asks whether the tool was used rather than how a decision was made. And it is not automated assessment, which concerns who or what does the marking: a machine-graded staged portfolio can be process-oriented, and a human-graded final essay need not be. The concept concerns which evidence is collected, over what period, and under what conditions it is produced.
The idea has moved from a proposition to a recurring recommendation across the knowledge base. It appears as a named priority in AIED's agenda for learner agency and Motivation, as the organizing research agenda for the retention and transfer problem in performance-versus-learning research, as "Black Box Assessment" in Brunnström and Palmqvist's account of self-regulated learning with AI, and as a request from practicing teacher educators who found their take-home written work could no longer evidence learning at all.
What process evidence looks like in practice
The most explicit design vocabulary comes from Thapa and Lewis, whose four commitments carry the argument into task design: reflective justification, staged task design, dialogic engagement, and evaluative transparency. In practice, students explain interpretive decisions rather than only presenting conclusions, submit work that develops across checkpoints, take part in oral or dialogic exchanges about their reasoning, and work against criteria they can apply themselves — the terrain of feedback literacy and Self-Assessment. Boysen's (2026) proposal for open learning practices translates the same logic into an audit trail: a shared learning plan set before the project begins, a learning analysis plan naming required and forbidden sources, and documented drafts and version histories, preserved through Track Changes and saved file versions, that expose the work process and explain what informed each revision. He frames this as self-regulation made visible.
Practitioner and framework accounts fill in the artifacts. Uden and Hwang's (2026) LEARN framework, aimed at problem-based learning Assessment, asks for process-focused rubrics, oral justifications, staged submissions, and reasoning traces with AI-use disclosure, and treats reflection through learning journals and think-aloud protocols as the metacognitive engine of the sequence. Moganadas and colleagues (2026) make AI-transparent, process-oriented assessment one of five researchable propositions, requiring documentation, verification, and reflective justification rather than prohibition or detector use. Padhy (2026) recommends capturing revision history, interaction patterns, and time-on-task, paired with transparent disclosure. Goldstein, Marae-Haj and Zidan (2026) report the same artifacts emerging from practice after a Trust crisis: prompting records, document version history, monitored group contributions, and reflective "journey journals" kept in physical notebooks — alongside the plain finding that take-home written work could no longer evidence learning.
A different axis from authenticity, and from marking
Holding the axes apart matters, because conflating them is how a redesign looks done while the exposure remains. Sharma (2026) makes the complementary distinction between authenticity and integrity-oriented design: authenticity asks whether a task mirrors worthwhile practice, while integrity-oriented design asks whether learners can justify their decisions and assume responsibility in relation to disciplinary standards, which is why the practices he offers — annotated decision trails, verification of GenAI-contributed claims, oral defense, draft differences with version history — are selected for the judgment they make visible rather than for their realism. Espino and Espino's (2026) ten-year review of business education reaches the same design conclusion from the literature side, finding the field converging on authentic, process-visible tasks that require interpretation, contextual application, and iterative refinement, and treating isolated tool adoption inside unchanged assessment structures as its most persistent failure.
The construct is also distinguishable from the question of how evidence is scored. Validity reasoning is where the two meet: Brunnström and Palmqvist (2026) argue that when a chatbot can produce a plausible examination answer, the submitted product becomes a weaker indicator of what the student learned — an assessment-validity problem rather than only an integrity problem. That is the same inference at risk in Yan, Greiff, Lodge and Gašević's performance-versus-learning argument, where performance gains that shrink once assistance is withdrawn show a gain living in the tool rather than in the learner, and where the proposed remedy is process measures such as retention and transfer tests rather than immediate task success.
Why detection cannot carry this load
Process-oriented assessment is usually proposed as the alternative to detection, and the evidence about detection explains why. The detection literature finds error rates, task-dependence, and patterned bias against non-native and disabled writers, and the most recent work lands on a procedural conclusion: a detector score can prompt closer review but is not a finding. Munoz and colleagues (2026) show what that costs in practice. Coding 1,162 generative-AI misconduct allegations, they found the strongest evidence types are generated by the investigation or by supervision, that process evidence such as drafts, supervision meetings, and presentations appears only where those practices already exist — so in an unredesigned assessment it is simply absent — and that detector output drew the lowest probative ratings of any category. The design consequence is ordering: process evidence has to exist before a case needs it, and Weidlich's (2026) reasoning about which inference is at risk makes the same point from the validity side, since attempts to restore assessment security through detection introduce construct-irrelevant variance rather than evidence about the learner.
The design and workload costs
The sources are unusually candid that this is expensive. Thapa and Lewis state the equity caution directly: reflective and dialogic tasks can privilege students who are more confident in academic self-articulation unless they are inclusively designed and scaffolded, and the relational, feedback-intensive work intensifies the emotional and professional labor of educators in large or resource-constrained settings — a contradiction with institutional expectations for scalability, standardization, and measurable performance. Sharma adds the interpretive version: requiring documented reasoning privileges Learners more fluent in reflective discourse, and judgment as evidence "remains relational and situated rather than mechanically verifiable."
The operational costs are itemized elsewhere. Boysen lists new technology, revised assignments and syllabi, teaching students unfamiliar skills, and extra material to collect and grade, warns that low-stakes credit for process can inflate grades, and concedes the model suits large skill-development projects far better than day-to-day knowledge-acquisition work; he also notes that process tracking is what some critics call AI surveillance. Goldstein, Marae-Haj and Zidan's participants reported that establishing process documentation took real work, that nine of thirteen funded advanced tools from their own pockets in a way that could deepen the digital divide, and that several fell back on supervised examinations and anticipated oral thesis defenses. The equity worry has an AI-specific edge in Brunnström and Palmqvist's finding that unguided tool use demands an interaction-management skill that is unevenly distributed, so AI "may be most beneficial to already advantaged students" — the same distributional risk that attends asking students to manage a documented process. Lopez-López, Bru-Cordero and Correa-Álvarez (2026) found that students who invested more independent study time judged AI use more strictly, and recommend testing whether disclosure templates, oral defenses, and process-based assessment reduce ethical uncertainty — a study the field still needs.
What remains unresolved
Thapa and Lewis's account is conceptual: no participants, no data collection, and no implementation trial, so the four principles are argued rather than tested, and the paper itself calls for empirical work on student experience, educator workload, and long-term viability. Feasibility is left undifferentiated across settings — transfer to large lectures, laboratories, studios, or clinical placements is not addressed. Lane's (2026) position paper notes the institutional ceiling: process-based assessment is blunted by grading and transcript structures that reward polished final artifacts, so classroom redesign alone does not change the incentive. Measurement is thin as well: Jin and colleagues' (2026) review of AI-literacy instruments calls for performance-based and process-oriented assessment and behavioral-trace measurement of prompting and verification, but reports that the existing instruments lean on self-report. No source here reports a controlled trial of a process-oriented redesign against a product-oriented one, which leaves the central claim — that process evidence is both harder to counterfeit and fairer to judge — argued from principle, practitioner report, and case-file evidence rather than demonstrated at scale.
Connected Concepts
- Assessment
- Authentic Assessment
- AI Detection
- Assessment Validity
- Evaluative Judgment
- Formative Assessment
- Summative Assessment
- Feedback
- Self-Assessment
- Oral Assessment
- E-Portfolio
- AI Use and Disclosure Statements
- Self-Regulated Learning
- Academic Integrity
- Equity
Connected Articles
- Preserving epistemic authenticity: process-oriented assessment in the age of generative AI — epistemic authenticity and the four design commitments behind process-oriented assessment (Thapa & Lewis 2026)
- Distinguishing performance gains from learning when using generative AI — performance is not learning: a research agenda built on process measures (Yan et al. 2026)
- AIED's Unfinished Mission: Centering Agency and Motivation in the Age of Effortless Bypass — amplifying process-based assessment as one of five AIED priorities (Lane 2026)
- AI-interaction literacy: reflections on how generative AI might be used to support self-regulated learning in higher education — Black Box Assessment and the learning trajectory in take-home examinations (Brunnström & Palmqvist 2026)
- The Safety Gap: Restoring Productive Struggle Through Pedagogically Aligned Generative AI — the safety gap and prioritizing process-based assessment in medical education (Wang & Shan 2026)
- Human-centered AI for teacher educators: Designing professional learning for critical AI literacy — teacher educators asking for AI-resistant, process-based tasks built on justification and reflection (Baran et al. 2026)
- Pioneering Teacher Educators Navigating AI Integration in Pre-Service Teacher Preparation: Strategies and Challenges — prompts, version history, monitored contributions, and journey journals after a trust crisis (Goldstein et al. 2026)
- Academic Integrity in the Age of AI: University Students' Study Practices and Ethical Judgments — study practices, ethical judgments, and the case for testing oral defenses and process evidence
- The LEARN Framework for Responsible Use of Generative AI in Education: A Neuroscience-Informed Model for Problem-Based Learning — the LEARN framework: process-focused rubrics, oral justifications, staged submissions (Uden & Hwang 2026)
- Generative AI as a Didactic-Pedagogical Mediator: Rethinking Human Roles and Pedagogical Design in Higher Education — AI-transparent process-oriented assessment as a researchable proposition (Moganadas et al. 2026)
- Is using artificial intelligence tools for academic work cheating? Student perceptions, ethics, and the impact — revision history, interaction patterns, and time-on-task with disclosure (Padhy 2026)
- Mapping the Integration of AI into Business Education: Insights from a Decade of Research — a decade of business-education research converging on process-visible tasks
- Measuring Artificial Intelligence Literacy: A Systematic Review of Instrument Development, Conceptual Foundations, and Psychometric Quality — the measurement agenda: performance-based and process-oriented assessment (Jin et al. 2026)
- A Proposal for Open Learning Practices in Response to Generative Artificial Intelligence — documented drafts and version histories as the auditable process record (Boysen 2026)
- Educational integrity in GenAI-augmented assessment: making judgment visible — annotated decision trails, oral defense, and draft differences selected for visible judgment
- How strong is the evidence in generative AI-related academic misconduct allegations? A mixed-methods analysis — what misconduct case files contain, and the process evidence they lack
- Which inference is at risk? Assessment validity reasoning and generative AI — which inference is at risk when assessment evidence is substituted
- Beyond Detection: Redesigning Authentic Assessment in an AI-Mediated World — authenticity must be redesigned, not policed
- From authentic products to authenticated processes: a systematic conceptual review of authentic assessment in AI-rich — authentic products and the processes that produced them
- Reconsidering the Use of Oral Exams and Assessments: An Old Way to Move Into a New Future — the oral exam as an AI-resistant format that makes reasoning live
Connected Resources
- Process FeedbackA free, process-based alternative to AI detection: it records how a piece of writing was produced and turns that into a report for a conversation about learning rather than a verdict.