π Research Article
Assessing students' DRIVE: A framework to evaluate learning through interactions with generative AI
Synthesis: Oliveira, Zednik, Bombaerts, Sadowski and Conijn (2025) propose DRIVE (Directive Reasoning Interaction and Visible Expertise), a conceptual framework and practice-oriented taxonomy for evaluating student learning from their interactions with generative AI chatbots in authentic, high-stakes classroom writing. Because the final essay is an increasingly ambiguous signal of learning when GenAI can produce indistinguishable text, the framework shifts assessment focus from the final product to the observable learning process β specifically how students steer the dialogue (Directive Reasoning Interaction, DRI) and how they make acquired course knowledge visible within it (Visible Expertise, VE). In a multi-methods analysis of 1,450 annotated GenAI interactions from 70 graded essays across three STEM philosophy and ethics courses, process-focused interaction-quality scores correlated strongly with traditional essay scores (r = 0.54), validating the approach, and the taxonomy revealed distinct interaction profiles β a "targeted improvement partnership" behind high essay scores, a "collaborative intellectual partnership" behind high interaction scores, and "basic information retrieval"/"passive task delegation" profiles behind below-average performance. The findings demonstrate how the choice of assessment method (output vs. process focus) shapes and rewards different types of student engagement with AI.
Key Findings
DRIVE defines two assessable components of learning. Directive Reasoning Interaction (DRI) captures how actively and purposefully a student steers the interaction with AI β strategic steering, critical questioning of AI output, and taking a leading role in the humanβmachine dialogue β grounded in heutagogy, the ICAP cognitive-engagement hierarchy, and "active human agency." Visible Expertise (VE) captures the extent to which a student makes acquired, course-specific knowledge visible in the dialogue, resonating with Harvard Project Zero's "Making Thinking Visible" and research linking insightful questioning to expertise. It is the combination of high DRI and high VE, with both defined against an educator's specific learning objectives, that makes the learning process and critical thinking assessable.
Process-focused assessment is empirically validated. Across 70 annotated essays, the GenAI interaction-quality score correlated strongly and positively with traditional essay scores (Pearson r = 0.54, 95% CI [0.34, 0.68], p < .001; Spearman rho = 0.54, p < .001), providing initial evidence that analyzing the interaction process is a valid approach to assessing learning in higher education and that simply using GenAI was not itself associated with essay performance (no significant essay-score difference between AI users and non-users).
Distinct GenAI interaction profiles emerge. High essay scores were linked to a "targeted improvement partnership" focused on systematic text refinement (Writing_Improve), critical engagement with AI content (Content_Critical), and relating concepts (Argument_Relate). High interaction scores were linked to a "collaborative intellectual partnership" centered on idea development β conceptual clarity, relating concepts, and bringing original ideas to the AI (Content_Idea). Below-average scores were associated with "basic information retrieval" (Content_Research, Content_Example) and "passive task delegation" (Writing_Instructions, Writing_Summarize, Writing_AutoComplete), where students outsourced cognitive work rather than collaborating.
Assessment method shapes and rewards GenAI usage. Confidence-interval overlap was high (21 of 22 classifications agreed), but paired comparisons revealed systematic differences: process-focused evaluation rewarded conceptualization-related work (Argument_ConceptualClarity, Content_Idea) and flexible prompting, while output-focused essay evaluation relatively preferred structured task specification (Writing_Instructions) and compensatory information-seeking (Content_Research). Traditional assessment can reinforce text optimization, whereas process-focused evaluation may reward an exploratory intellectual partnership with AI.
A practice-oriented taxonomy operationalizes the framework. A 35-subcategory taxonomy (13 Writing, 10 Content, 12 Argument) was developed through iterative Design-Based Research by teachers and teaching assistants, with moderate average inter-rater agreement (Cohen's ΞΊ = 0.44). Writing interactions dominated (41.3%), followed by Content (28.7%) and Argument (22.3%). The taxonomy gives educators a concrete, tool-agnostic analytical instrument for detecting evidence of learning in student-GenAI dialogue.
The DRIVE framework
The emergence of generative AI in higher education has fundamentally disrupted traditional Assessment, especially in academic writing, where AI can produce text increasingly indistinguishable from human work. When the final product no longer provides a clear signal of a student's knowledge or skills, conventional output-focused assessment becomes an unreliable measure of learning, and the focus must shift to the learning process itself. Analyzing the dialogue between a student and a GenAI system provides a more transparent record of that process, allowing educators to observe both how students actively steer the AI (evidence of directive reasoning) and how they make acquired domain-specific knowledge visible through how they deploy it in these interactions.
DRIVE was developed using an approach aligned with Design-Based Research principles, working in a setting with high ecological validity: university courses where students used GenAI to complete graded assignments and submitted their interaction logs as a formally assessed component. The initial impetus came from teachers who, during early ungraded experiments, recognized the need for a structured method to move beyond intuitive judgments when analyzing student learning in chatbot interaction logs. The framework is deliberately flexible: what counts as high-quality DRI and VE is not universal but is defined by an educator's specific learning objectives, positioning the educator's pedagogical goals as the ultimate benchmark for what counts as meaningful evidence of learning in a student-GenAI interaction.
DRI is grounded in long-standing educational theory β heutagogy (self-determined learning), constructivism, and the ICAP framework, which differentiates cognitive engagement from passive reception through to constructive and interactive engagement. A high DRI profile reflects "active human agency," serving as a cognitive safeguard against automation bias and the skill atrophy associated with Cognitive Offloading, through which a person reduces cognitive effort by delegating a task to AI. VE, in turn, is grounded in the idea that thinking must be made observable to be understood, directed, and assessed; a prompt can signal the nature of the asker's knowledge structure, since more knowledgeable individuals are better able to identify gaps in information and formulate the specific questions needed to fill them. This visibility offers a window into the student's learning process and is essential for accountability and Trust in the classroom.
The interaction taxonomy
To operationalize DRI and VE and systematically analyze interactions, the authors developed a detailed interaction taxonomy with three main categories aligned with argumentative writing theory (Toulmin; Wingate): Writing (mechanical and structural aspects of composition), Content (knowledge construction and understanding, with emphasis on course-specific material), and Argument (logical and analytical aspects of writing). The taxonomy was built through an iterative, practitioner-led process combining a top-down component based on pedagogical goals with a bottom-up analysis of actual student-GenAI interaction logs, yielding 35 subcategories. Annotators classified each student prompt with the best-fitting taxonomy item(s); interactions fitting more than one category were labeled "Mixed."
The taxonomy serves as the analytical tool for classifying the specific prompts that, viewed holistically, provide evidence of a student's DRI and VE profile. It is designed to be tool-agnostic: the evaluation focuses on the quality and nature of the students' prompts as indicators of their learning and agency, independent of the specific capabilities of the AI model. This addresses a core challenge in Assessment with GenAI β the need to distinguish between general AI Literacy and domain-specific learning β by evaluating whether the content of a student's prompts reveals their unique, acquired knowledge of course material, which is critical for Authentic Assessment.
Study design and method
The study was conducted at a STEM university across three Bachelor's or Master's level courses on philosophy and ethics during the 2023β2024 and 2024β2025 academic years. Of 445 enrolled students, 103 (23.2%) chose to use GenAI under the condition that they submit their interaction logs for assessment; a subset of 70 essays and their corresponding interaction logs were annotated, yielding 1,450 annotated prompts. ChatGPT was the most common tool (68.6%). Student performance was assessed using two measures: a traditional output-focused essay score and a novel process-focused GenAI interaction-quality score, both standardized into z-scores within each course. A grading policy specified that the final grade for GenAI users was a weighted average (2/3 essay, 1/3 interaction).
Inter-grader reliability was high, with low weighted mean absolute deviation (Weighted MADs = 0.035 for essays, 0.209 for interactions). Notably, teachers observed a marked decrease in students willingly adopting GenAI after they began formally grading interactions β students may have perceived the traditional essay-only path as less demanding or lower risk. The interaction-quality rubric integrated three criteria β "AI for Writing," "AI for Argumentation," and "AI for Course Content" β aligned with the taxonomy and linking to the DRIVE framework by focusing on agentic cognitive engagement (DRI) and visible knowledge integration (VE). This design situates the work within the broader program of Learning Analytics, using process data (interaction logs) as evidence of learning, and connects to Educational Measurement by demonstrating how a novel process-focused measure can be validated against an established learning outcome.
Interaction profiles and what they reward
The taxonomy descriptives revealed that most interactions clustered around average mastery β 77% of essay classifications and 82% of interaction classifications were average β underscoring that only a few specific prompting strategies are connected with very high and very low quality on each measure. High essay mastery was characterized by a "targeted improvement partnership": students engaged GenAI as a tool to systematically refine their own input work through critical evaluation and conceptual integration rather than seeking comprehensive assistance. High interaction mastery was characterized by a "collaborative intellectual partnership": students engaged GenAI as an intellectual collaborator, bringing original, well-motivated ideas (Content_Idea), seeking conceptual clarity, and critically evaluating AI outputs. By contrast, below-average essay mastery showed a "basic information retrieval" profile (asking the AI to define ideas, find related concepts, or supply examples), and below-average interaction mastery showed a "passive task delegation" profile (specifying tasks by copy-pasting assignment instructions, asking the AI to summarize or auto-complete) β behaviors that outsource cognitive work and show minimal DRI and VE.
The systematic, though modest, differences between assessment methods highlight that each approach offers a particular perspective from which to evaluate student work. Because the focus of the assessment β grading the process (interaction log) or the product (essay) β can interact with the writing task to shape what types of GenAI interactions are ultimately recognized and rewarded, the findings connect to ongoing discussions in Formative Assessment and assessment in technology-enhanced learning. The framework also offers a granular lens on the "performance versus learning" paradox, distinguishing interactions that primarily leverage AI for a high-quality output from those showing the student actively thinking with and through the tool. This is directly relevant to Academic Integrity conversations: by making the student's intellectual contribution visible within the interaction, process-focused assessment provides a more transparent basis for attributing work than the ambiguous final product alone.
Implications
For educators in writing-intensive courses where GenAI use is permitted, the findings suggest that combining traditional writing assessment with interaction-log evaluation captures complementary aspects of student work: traditional assessment identified strengths in text refinement and conceptual integration, while interaction-log evaluation revealed critical-thinking processes and sophisticated AI-collaboration strategies not always evident in the final text. Teachers designing assessments for AI-integrated writing should consider how the evaluation focus shapes and rewards different GenAI-usage patterns, and AI-related grading rubrics should distinguish between different types of interaction patterns based on course learning goals, recognizing that exploratory, conceptual development is rewarded by process-focused evaluation even when it does not directly raise output quality. Because the adoption rate dropped when interactions became a graded component, understanding student perspectives about process-focused assessment is valuable before implementation. While automated classification may eventually assist with log evaluation, human oversight remains essential for accurately assessing sophisticated collaboration β a caution relevant to Learning Analytics pipelines. The framework positions the educator's pedagogical goals as the benchmark for meaningful learning evidence, offering a practical design for Authentic Assessment that captures learning in AI-integrated classrooms.
Connected Concepts
- Assessment
- Learning Analytics
- Authentic Assessment
- Generative AI
- Cognitive Offloading
- Self Regulated Learning
- Higher Ed
- Trust
- Educational Measurement
- Formative Assessment
- AI Literacy
- Academic Integrity
Connected Articles
- Aaiwa AI Authentic Assessment Metacognition 2026 β Authentic AI-assisted assessment and metacognition
- Agency Gap AI Writing β Student agency in AI-assisted writing
- LLM Formative Feedback Systematic Review 2026 β LLM-generated formative feedback
Citation
Oliveira, M., Zednik, C., Bombaerts, G., Sadowski, B., & Conijn, R. (2025). Assessing students' DRIVE: A framework to evaluate learning through interactions with generative AI. Computers and Education: AI, 9, 100497.