Research Article
Characterizing Students' LLM Usage Behaviors and Their Association with Learning in Critical Thinking Tasks
Synthesis: This study extends prior work on student LLM use by analyzing data from two offerings of a research-oriented course where students learn to read, reason about, and critique academic papers — a setting that moves beyond the Problem Solving domains that dominate existing research. Crucially, students had no restrictions on LLM usage, providing ecological validity.
Key contributions:
- Refined bottom-up categorization of LLM usage types in academic critical thinking, cross-labeled by the extent of student initiative — from passive (copy-pasting text for summaries) to active (using LLM as a Socratic dialogue partner for Scaffolding critical thinking with generative AI: Design principles for integrating large language models in higher).
- Learning outcome analysis linking usage frequency and type to performance on three midterm exams. This addresses the core question raised by Distinguishing performance gains from learning when using generative AI: do specific LLM usage patterns help or hinder actual learning?
- The student initiative dimension is particularly valuable for understanding AI Literacy development — it maps onto the distinction between using AI as a crutch vs. as a cognitive tool, directly relevant to Scaffolding design.
This work complements Not All Students Engage Alike: Multi-Institution Patterns in GenAI Tutor Use by shifting focus from tutoring to student-initiated LLM use in authentic academic tasks. The EDM 2026 acceptance places it within the Learning Analytics community's growing interest in modeling AI-augmented learning behaviors. Findings also inform Educational Development strategies for guiding student AI use.
What this means for practice
- Instructors. Front-load explicit guidance on LLM use in the first weeks of a course: students who reported no LLM use outperformed users on Midterm 1 with a large effect size, and only later in the term did the LLM group close the gap by roughly 10%.
- Instructors. Teach the difference between student-driven and LLM-driven prompting rather than banning or permitting tools wholesale, because students whose use was mostly student-driven scored about 10% higher on Midterm 1 than those relying on LLM-driven support.
- Instructors. Track how much of the work is delegated, not just whether a tool was opened: the 7 High-Reliance students (LLM use in more than 50% of submissions) trailed the 16 Low-Reliance students (5–33% of submissions) by 7–8% on all three midterms.
- Learners. Use the LLM to interrogate your own reading of a paper — the student-driven patterns in this course, such as soliciting counterarguments, were associated with better exam performance than asking the model to generate content.
Limitations
- The data cover 68 students across two offerings of a single research-oriented course (37 and 31), so the usage patterns are specific to that course and setting.
- Usage frequency and type come from students' self-reported weekly assignments rather than logged interactions, so the categories rest on what students chose to record.
- Subgroup comparisons between High-Reliance and Low-Reliance students were not tested statistically because the groups were highly imbalanced (7 vs 16), so those differences are descriptive trends.
- Students self-selected into LLM use with no restrictions and no control group, so the midterm gap cannot be read as an effect of LLM use.
Citation
Park, M., Orozco Vasquez, I., & Conati, C. (2026). Characterizing students' LLM usage behaviors and their association with learning in critical thinking tasks. In Proceedings of EDM 2026.