On this page

Synthesis: This design science research (DSR) develops a rule-guided LLM tutoring system for primary-school mathematical word problems, addressing the inconsistency and pedagogical opacity of unconstrained intelligent tutors in a procedural domain where correctness is governed by deterministic rules. The artifact formalizes a distinction between rule-guided scaffolding, governed by a three-layer architecture (diagnosis → intent selection → constrained response generation), and ad-hoc scaffolding, where helpful moves are difficult to audit and replicate. The system constrains stochastic variability through explicit Guardrails while embedding Scaffolding theory, evaluated through a staged DSR approach spanning persona-based simulated dialogues and a real classroom pilot. Findings show rule-guided scaffolding improves interactional consistency, reduces premature answer-giving, and sustains cognitive engagement — while revealing interactional complexities that only authentic classrooms expose.

Key Findings

  • Rule-guided scaffolding outperformed ad-hoc scaffolding on interactional consistency: in 36 of 40 (90%) real-student sessions both the tutor and students followed the intended turn-taking script, and the tutor guided key concepts/rules in the same 36 sessions, even under fragmented, ambiguous classroom input.
  • The three-layer architecture reduced premature answer-giving and early closure via anti-spoiler boundaries and a completion-only goodbye gate; the tutor's diagnostic layer identified and responded to clear student deviations in all sessions where they occurred, validating the granularity of its assessment codes.
  • Simulation surfaced architectural failure modes with high recall — arithmetic verification misjudgment, over-explanation, answer-giving boundary violations, and premature goodbye were all identified and remediated through an "evidence–diagnosis–revision–retest" cycle using an LLM-as-Student, dual-model configuration (DeepSeek as student, ChatGPT as tutor).
  • The classroom pilot with 40 Grade 5 students (mean age 10.8 years) revealed complexities not captured in simulation: attentional fragility requiring pause/resume, engagement costs of pedagogical verbosity, trust sensitivity to arithmetic disagreements, and need for explicit locking of high-risk procedural rules (units, rounding, billing conventions).
  • Even a highly structured prompt cannot fully eliminate stochastic variability — the tutor produced incorrect information (arithmetic hallucinations) in 4 of 40 sessions (10%), motivating strengthened uncertainty guardrails and an "epistemic humility" design in which all computation is returned to the student.
  • Simulation and classroom validation are complementary, not interchangeable: every Phase 1 failure mode reappeared in Phase 2, yet simulation systematically underestimates interactional failure modes (fragmented inputs, off-task behavior, trust sensitivity), warning against "ecological overconfidence."

Implications for Practice

  • For designers of LLM tutors: effective tutoring is less about maximizing the model's generative capacity than about strategically constraining it within theoretically grounded frameworks — externalize pedagogical decisions into auditable diagnosis → intent → response layers rather than relying on unconstrained inference.
  • For mathematics and procedural-domain applications: embed explicit Guardrails such as unit/rule locking, numerical correctness gates, and anti-spoiler boundaries; consider removing arithmetic from the tutor's purview entirely and returning all computation to the student (epistemic humility).
  • For evaluation practice: use staged design science evaluation — simulated persona stress-tests for diagnostic precision plus authentic classroom pilots for ecological validity — since strong simulation performance does not guarantee classroom readiness.
  • For developers managing cognitive load and engagement: tighten brevity and anti-repetition output constraints, add attentional scaffolds (e.g., pause-close for resuming interrupted work), and communicate transparent closure criteria to reduce premature termination.

Connected Concepts

Connected Articles

Citation

Taming the black box: Design principles for rule-integrated LLM tutoring systems in primary school mathematical problem solving — Looi, C.-K., Liu, Z., & Sun, D. (2026). Computers and Education: Artificial Intelligence, 10, 100586.

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.