On this page

Synthesis: This large-scale quasi-experiment (N = 313 sophomores, four authentic classes, four weeks) evaluated DBagent — a domain-specific LLM-based educational agent for an undergraduate database course with tool use, memory, and goal-directed reasoning. The agent-enriched environment significantly improved learning achievement, but lag sequential analysis of interaction logs revealed a distinctive cognitive profile: high-frequency lower-order engagement (Remember/Understand, ~54.5%) driven by psychological safety, organized around a "Query-Evaluation-Query" verification loop — with only 3.92% of interactions reaching higher-order cognition. SEM confirmed positive perceptions sustain engagement via satisfaction.

Key Findings

  1. Improved learning achievement. The agent-powered context significantly outperformed traditional instruction; the top experimental class beat the control on both tasks (Z = 3.49, Z = 5.15, p < .001), with the largest effect on complex Problem Solving (Task 2).
  2. A shift from social inhibition to psychological safety. Student-agent interactions were dominated by lower-order cognitive activity (~54.5%) because the agent provided a judgment-free environment that encouraged Help-Seeking — interpreted as psychological safety rather than mere dependency.
  3. The "Query-Evaluation-Query" verification loop. LSA identified a significant QR-EA-QR loop: students offload recall/understanding to the agent, then transition into Evaluation of its output — an "offload-evaluate cycle" distinct from the linear confusion-to-understanding path of human-instructor interaction.
  4. But lower-order lock-in. Strong self-transition loops within lower-order states (Understand z = 49.08; Application z = 51.48) show learners get "locked" in routine processing, with only 3.92% reaching higher-order cognition — attributed to the agent's unwavering compliance lacking pedagogical friction.
  5. Perceptions drive engagement via satisfaction. SEM confirmed learners' positive perceptions of the agent promoted sustained engagement through the mediating role of satisfaction.
  6. The "prompt engineering gap." Efficacy was moderated by domain-specific digital readiness — a Geoscience-major class underperformed CS cohorts on Task 2, suggesting non-technical students need targeted Scaffolding to bridge the Prompt Engineering gap.

What this means for practice

  • Instructors. Do not read the achievement gain as evidence of higher-order learning: only 3.92% of student–agent interactions reached higher-order cognition, with strong self-transition loops in lower-order states (Understand z = 49.08; Application z = 51.48), so track cognitive engagement alongside test scores.
  • Instructors. Build verification and critical-evaluation scaffolds around the "Query-Evaluation-Query" loop that lag sequential analysis identified — students already move from offloading recall to evaluating agent output, and Cognitive Diagnosis should capture what they actually process, making that evaluation deliberate instead of incidental.
  • Designers. Design productive struggle into the agent rather than unconditional compliance, because the agent's unwavering helpfulness produced lower-order lock-in and the psychological-safety advantage carried a Cognitive Offloading risk; Self-Regulated Learning must be deliberately supported rather than assumed, and the same argument runs through Measuring Whether LLM Tutors Teach or Solve: A Diagnostic for Educational Impact and Productive Failure.
  • Designers. Add targeted Scaffolding for non-technical majors: a Geoscience-major class underperformed the CS cohorts on the complex Task 2, implicating a Prompt Engineering gap created by differing domain-specific digital readiness.
  • Researchers. Test the psychological-safety reading directly: the high lower-order share (~54.52%) is interpreted as safety rather than dependency, but the design measured interaction logs, not the students' reasons for asking.

Limitations

  • The quasi-experiment used four intact classes (three experimental, one control) with no random assignment, so class-level differences cannot be fully separated from the intervention.
  • The study covers one undergraduate database course over four weeks with two open-ended tasks, so the achievement evidence is short-term and single-course, and two experimental classes showed no significant gain on Task 1.
  • Cognitive-engagement findings come from lag sequential analysis of interaction logs, and the psychological-safety explanation is inferred from those sequences rather than measured.
  • Efficacy varied by cohort — a Geoscience-major class underperformed the CS cohorts on Task 2 — so the pooled improvement masks subgroup differences driven by domain-specific digital readiness.

Citation

Li, X., Liu, Z., Jiang, S., Chen, J., & Chen, W. (2026). The impact of an LLM-based educational agent on learning achievement, cognitive dynamics, and student perceptions in computer science education. International Journal of STEM Education, 13, 51.

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.