Research Article
The Impact of an LLM-Based Educational Agent on Learning Achievement, Cognitive Dynamics, and Student Perceptions in Computer Science Education
Synthesis: This large-scale quasi-experiment (N = 313 sophomores, four authentic classes, four weeks) evaluated DBagent — a domain-specific LLM-based educational agent for an undergraduate database course with tool use, memory, and goal-directed reasoning. The agent-enriched environment significantly improved learning achievement, but lag sequential analysis of interaction logs revealed a distinctive cognitive profile: high-frequency lower-order engagement (Remember/Understand, ~54.5%) driven by psychological safety, organized around a "Query-Evaluation-Query" verification loop — with only 3.92% of interactions reaching higher-order cognition. SEM confirmed positive perceptions sustain engagement via satisfaction.
Key Findings
- Improved learning achievement. The agent-powered context significantly outperformed traditional instruction; the top experimental class beat the control on both tasks (Z = 3.49, Z = 5.15, p < .001), with the largest effect on complex Problem Solving (Task 2).
- A shift from social inhibition to psychological safety. Student-agent interactions were dominated by lower-order cognitive activity (~54.5%) because the agent provided a judgment-free environment that encouraged Help-Seeking — interpreted as psychological safety rather than mere dependency.
- The "Query-Evaluation-Query" verification loop. LSA identified a significant QR-EA-QR loop: students offload recall/understanding to the agent, then transition into Evaluation of its output — an "offload-evaluate cycle" distinct from the linear confusion-to-understanding path of human-instructor interaction.
- But lower-order lock-in. Strong self-transition loops within lower-order states (Understand z = 49.08; Application z = 51.48) show learners get "locked" in routine processing, with only 3.92% reaching higher-order cognition — attributed to the agent's unwavering compliance lacking pedagogical friction.
- Perceptions drive engagement via satisfaction. SEM confirmed learners' positive perceptions of the agent promoted sustained engagement through the mediating role of satisfaction.
- The "prompt engineering gap." Efficacy was moderated by domain-specific digital readiness — a Geoscience-major class underperformed CS cohorts on Task 2, suggesting non-technical students need targeted Scaffolding to bridge the Prompt Engineering gap.
What this means for practice
- Instructors. Do not read the achievement gain as evidence of higher-order learning: only 3.92% of student–agent interactions reached higher-order cognition, with strong self-transition loops in lower-order states (Understand z = 49.08; Application z = 51.48), so track cognitive engagement alongside test scores.
- Instructors. Build verification and critical-evaluation scaffolds around the "Query-Evaluation-Query" loop that lag sequential analysis identified — students already move from offloading recall to evaluating agent output, and Cognitive Diagnosis should capture what they actually process, making that evaluation deliberate instead of incidental.
- Designers. Design productive struggle into the agent rather than unconditional compliance, because the agent's unwavering helpfulness produced lower-order lock-in and the psychological-safety advantage carried a Cognitive Offloading risk; Self-Regulated Learning must be deliberately supported rather than assumed, and the same argument runs through Measuring Whether LLM Tutors Teach or Solve: A Diagnostic for Educational Impact and Productive Failure.
- Designers. Add targeted Scaffolding for non-technical majors: a Geoscience-major class underperformed the CS cohorts on the complex Task 2, implicating a Prompt Engineering gap created by differing domain-specific digital readiness.
- Researchers. Test the psychological-safety reading directly: the high lower-order share (~54.52%) is interpreted as safety rather than dependency, but the design measured interaction logs, not the students' reasons for asking.
Limitations
- The quasi-experiment used four intact classes (three experimental, one control) with no random assignment, so class-level differences cannot be fully separated from the intervention.
- The study covers one undergraduate database course over four weeks with two open-ended tasks, so the achievement evidence is short-term and single-course, and two experimental classes showed no significant gain on Task 1.
- Cognitive-engagement findings come from lag sequential analysis of interaction logs, and the psychological-safety explanation is inferred from those sequences rather than measured.
- Efficacy varied by cohort — a Geoscience-major class underperformed the CS cohorts on Task 2 — so the pooled improvement masks subgroup differences driven by domain-specific digital readiness.
Citation
Li, X., Liu, Z., Jiang, S., Chen, J., & Chen, W. (2026). The impact of an LLM-based educational agent on learning achievement, cognitive dynamics, and student perceptions in computer science education. International Journal of STEM Education, 13, 51.