Research Article
ARPG+: a simulation-based study of real-time coaching for educational LLM prompting
Synthesis: Ye and colleagues present ARPG+, a real-time coaching system that teaches students how to prompt large language models (LLMs) effectively, grounding its design in cognitive load theory and the zone of proximal development. The system senses when learners struggle, delivers calibrated just-in-time interventions, and fades support as skills develop, tracking learner capability with uncertainty quantification and diagnosing prompt quality across six dimensions. Evaluation with LLM-based simulated learners showed prompt quality increases 143% beyond unguided practice and independence reaches 91% of final interactions versus 59% under fixed support, generalizing to other domains without retraining. The authors are explicit that all results are Simulation-based and that classroom validation is the necessary next step.
Core Finding
Prompting is a learnable, transferable metacognitive skill, and principled real-time coaching can improve prompt quality, accelerate learning, prevent cognitive overload, and foster durable autonomy — but effective Scaffolding must be able to stop helping, fading support as competence emerges to avoid dependency and metacognitive laziness. The design directly responds to the "autonomy paradox" in which generative-AI support that never withdraws hollows out the self-AI Regulation in Education that makes learning durable.
Operationalizing Learning Theory
ARPG+ converts abstract pedagogical constructs into calibratable, real-time decision variables. Cognitive load theory is instantiated through proxies for intrinsic, extraneous, and germane load aggregated into a bounded overload-risk signal; the zone of proximal development becomes a dynamic ability estimate with an uncertainty-aware moving boundary; and a normalized Struggle Index quantifies interactional friction from editing and pausing behaviors. Prompt quality is represented as a six-dimensional vector — structural completeness, semantic clarity, contextual grounding, task specificity, constraint explicitness, and output conventions — enabling fine-grained diagnosis and longitudinal tracking rather than holistic scoring.
Dual-Process Architecture and Scaffolding
A lightweight-deep dual architecture ensures fast responsiveness for routine interactions (50ms fast path) while reserving richer analysis for critical moments (300ms deep path). Coaching is cast as a constrained sequential decision process in which an information-theoretic selector optimizes Feedback content and granularity to maximize expected uncertainty reduction while bounding cognitive overload risk. A dynamic scaffolding-density schedule with exponential decay and periodic skill probes prevents pseudo-mastery: ablations show removing scaffolding drops independence from 0.915 to 0.662, and removing reinforcement primarily harms retention.
Simulated Evidence and Limitations
Across simulated learners, ARPG+ improved prompt quality by 143% beyond unguided practice, achieved 91% independence in late-session interactions versus 59% under fixed support and 71% under linear decay, and generalized across five additional domains (code writing, data analysis, creative design, academic writing, business communication) with 93.8% average retention. The authors are careful to frame this as system feasibility and simulation-based performance, not educational effectiveness, noting that simulated agents lack the affective, motivational, and interpersonal dynamics of real students and that equity-relevant dimensions (first-generation status, second-language learning, neurodivergence, cultural variation in Help-Seeking) remain unaddressed. A three-phase classroom validation agenda is laid out.
Relevance to the knowledge base
This paper advances the knowledge base's understanding of Prompt Engineering as a teachable skill rather than a mere technique, and its treatment of fading support directly addresses the knowledge base's concerns about Cognitive Offloading and AI-driven autonomy erosion. By operationalizing Self-Regulated Learning, Metacognition, and Learning Design in a real-time coaching loop, it demonstrates how Human AI Collaboration can be engineered to build rather than erode learner Learner Agency. Its explicit honesty about simulation-based limits is a model for evaluating Generative AI learning tools.
What this means for practice
- Designers. Build fading into the coach from the start: gate support removal on demonstrated competence and add periodic skill probes, because removing the scaffolding component dropped independence from 0.915 to 0.662 in ablation while removing reinforcement mainly harmed retention.
- Designers. Diagnose prompt quality across separate dimensions rather than scoring prompts holistically, so feedback can target the specific gap — structural completeness, semantic clarity, contextual grounding, task specificity, constraint explicitness, output conventions.
- Instructors. Coach prompting as a reasoning skill in the flow of work instead of handing out static templates: students reached a final quality of 7.82 under ARPG+ versus 5.95 under templates and 4.52 with no assistance.
- Researchers. Promote autonomy and transfer to first-class outcomes: judge tools on independence and transfer alongside output quality, since evidence on metacognitive laziness shows quality gains can coexist with hollowed-out regulation.
Limitations
- All results come from 1,000 LLM-based simulated learners (three Gaussian cognitive profiles in a 1:2:1 ratio, 20 turns each across five domains), not real students; the authors state the work is evidence of system feasibility and simulation-based performance, not educational effectiveness.
- Both sides of the simulation are LLM-generated — Doubao_lite drafts prompts as the behavior generator and Deepseek_v3 scores them as an independent evaluator — and the authors note the smooth trajectories and tight confidence bands are "partly a property of the simulation" rather than of the system alone.
- Simulated agents cannot report whether hints feel supportive or intrusive, and the study tests no instructor adoption, curriculum fit, time-on-task constraints, or coexistence with other tools, leaving learner experience and classroom integration outside the evidence base.
- Equity-relevant differences (first-generation status, second-language learning, neurodivergence, socioeconomic background, prior AI exposure, cultural variation in Help-Seeking) are not represented by the three profiles, and hyperparameters were frozen from a 200-learner held-out calibration set rather than validated against human data.
Connected Concepts
- Prompt Engineering
- Large Language Models (LLMs)
- Metacognition
- Cognitive Offloading
- Self-Regulated Learning
- Generative AI
- Human AI Collaboration
- Learning Design
- Learner Agency
Connected Articles
- Prompting for Teachability: Designing Novice Personas in LLMs for Learning by Teaching Contexts
- Same Question, Different Answer? Measuring and Mitigating Prompt Privilege for Equitable AI Access
- Think First, ChatGPT Later: Guiding Human–AI Collaboration for Learning Gains in Independent Human Creativity
- The Impact of Large Language Models on Programming Education and Student Learning Outcomes
Citation
Ye, P.-G., Mo, K., Long, Y., Liu, M., Sang, H., & Zheng, J. (2026). ARPG+: a simulation-based study of real-time coaching for educational LLM prompting. International Journal of Educational Technology in Higher Education.