Synthesis: An LLM-based interactive module teaches K-12 students prompting literacy through scenario-based deliberate practice with an AI auto-grader providing immediate, detailed feedback. Deployed across 11 secondary classrooms in two iterations, the module improved students' prompting skills (particularly embedding background context) and confidence in using AI for learning. The study also validates an AI-based auto-grader achieving 0.92 average accuracy and identifies True/False + open-ended questions as more effective than MCQs for assessing prompting literacy.
Study Design
Xiao et al. designed and deployed a web-based interactive instructional system to teach prompting literacy to secondary-education students. The module was grounded in two learning sciences principles: learning-by-doing and elaborated immediate feedback. Students practiced prompt writing in three hypothetical learning scenarios (biology, geography, math), each paired with a unique instructional activity (extending knowledge, quiz preparation, homework struggle).
After a student writes a prompt, an LLM-based auto-grader (GPT-4o) evaluates it across preset dimensions and delivers immediate, detailed feedback. The pipeline mirrors authentic AI chatbot interaction: write prompt β receive AI response β get graded feedback.
Two iterations were conducted across 11 secondary classrooms:
Study 1 (June 2024): 111 students, 6 classrooms in East AsiaStudy 2: Assessment iteration follow-up with similar populationThe Auto-Grader
The AI auto-grader achieved 0.92 average accuracy across dimensions when grading student-written prompts, using human labels as ground truth:
| Dimension | Accuracy |
|---|
| Relevance | 0.98 |
| Background/Context | 0.96 |
| Conciseness | 0.93 |
| Elaboration | 0.90 |
| No Direct Answer | 0.88 |
| Clarity of Purpose | 0.85 |
The lowest accuracy (Purpose, 0.85) stemmed from the auto-grader over-generating keywords or conflating Purpose with No Direct Answer criteria. The auto-grader tended to weigh heavily on some keywords while ignoring others β a known limitation of LLM-based grading.
Key Findings
Prompting Skill Improvement
Students improved significantly at embedding background/context information in prompts (McNemar test, p = .039 from Q1 to Q3)Students performed well on Relevance, Conciseness, and Purpose even in the first question (ceiling effects)Prior AI usage frequency was positively correlated with initial prompt quality (r = 0.27, p = .017), suggesting an equity concernConfidence and Perception
Self-reported confidence in using AI for learning increased by 10.4% (p < .001)87% of students reported learning AI-related knowledge (how to use AI for learning, how to ask effective questions, AI's capabilities)Students valued: direct AI interaction, scenario-based design, immediate comprehensive feedback, and visual elementsAssessment Design Lessons
MCQs suffered from ceiling effects β students could identify good prompts conceptually but couldn't write them effectivelyTrue/False + open-ended questions demonstrated better item difficulty and discrimination than MCQsNone of the original MCQ items fell into the desired difficulty range [0.3, 0.7], while 60% of OE and 30% of TF questions didChallenges Identified
Productive struggles: difficulty writing effective prompts (the core skill being taught)Extraneous load: slow AI response times, login issues, limited typing skills (22 students reported this)Scenario variety: some students wanted non-STEM scenariosLLM response latency disrupted the practice flowDesign Implications
The study demonstrates that Prompt Engineering can be taught effectively to K-12 students through structured practice with automated feedback. Key design principles:
1. Scenario-based deliberate practice with authentic AI interaction
2. Immediate, dimension-level feedback powered by LLM auto-grading
3. Assessment aligned to competency β open-ended + T/F outperform MCQs for higher-order prompting skills
4. Addressing the digital divide β prior AI access correlates with initial performance, underscoring the need for in-school prompting literacy instruction
Connected Concepts
AI EducationAI LiteracyAutomated GradingK 12LLMPrompt EngineeringStudent ExperienceRAGConnected Articles
Academiclaw Student Agent Benchmark β AcademiClaw: When Students Set Challenges for AI AgentsAccess Not Enough AI Tutoring 2026 β Access is Not Enough: Human Support Improves Engagement with AI TutoringAdapt Adaptive Lesson Plan Transformer β AdaPT: Adaptive Lesson Plan Transformer for Cross-Regional and Differentiated InstructionAffective Text Wearable Student Health β A Formative Study of Brief Affective Text as a Complement to Wearable Sensing for Longitudinal Student Health MonitoringAgency Gap AI Writing β The agency gap in AI-supported writing: how reactive and proactive agent designs shape multimodal reasoningAgent Voice Accents K12 Group Learning β Exploring How Agent Voice Accents Shape Human-AI Collaboration in K-12 Group LearningAgentic AI Education Scoping Review β Agentic AI in Education: A Scoping Review of Research Landscape, Capabilities, and the Frontier Agent ParadigmAgentic Literacy Debt β Agentic Literacy Debt: A Structural Problem the AI Literacy Field Has Not Yet NamedAgentic Workflows Education β Agentic Workflows in EducationAgents That Teach Incidental Learning β Agents That Teach: Designing Incidental Learning Back into AI-Assisted Software DevelopmentAgreement Not Quality LLM Coding Verification β Agreement Is Not Quality: Blind Expert Verification of Human and LLM Qualitative Coding When Human Consensus Is Not G...AI Adoption Training Public Sector β The Main Barrier to AI Adoption in the Public Sector is Lack of TrainingAI Adult Learning Guidelines Dis2026 β Guidelines for Designing AI Technologies to Support Adult LearningAI Agents Peer Learning Discourse β When AI Agents Teach Each Other: Discourse Patterns Resembling Peer Learning in the Moltbook CommunityAI Assessment Human Tutors β AI-Driven Assessment of Human Tutors: Linking Training Performance to Real-Life PracticeAI Assessment Scale Reform β A bit of chaos and madness": The AI Assessment Scale and the work of assessment reformAI Assistance Discretionary Feedback β AI Assistance for Discretionary Work: Increasing Feedback Provision in Higher EducationAI Assisted Learning Modes Eeg β An exploratory behavioral and electroencephalographic study of artificial intelligence-assisted learning modes in hig...AI Assisted Se Curriculum Syllabus Analysis 2026 β Mapping the Emerging Curriculum for AI-Assisted Software Engineering via Syllabus AnalysisAI Availability Student Motivation β Why Put in This Much Effort?": How AI Availability Shapes Studentsβ Motivation in Introductory ProgrammingAI Campus Wellbeing Tools β AI-Driven Tools for Enhancing Campus Well-being: Prevention and InterventionAI Changing Teaching Workflows β How AI Is Changing Teaching WorkflowsAI Education Global Capacity β What AI in Education Needs Next: Lessons from Youth Leaders Across Five CountriesAI Enabled Serious Games β AI-Enabled Serious Games: Integrating Intelligence and Adaptivity in Training SystemsAI Engineering Education Balancing Act β Using AI in engineering education: a balancing act, driven by clear purposeCitation
Xiao, R., Hou, X., Tseng, Y.-J., Nieu, H., Liao, G., Stamper, J., & Koedinger, K. R. (2026). Learning to Use AI for Learning: Teaching Responsible Use of AI Chatbot to K-12 Students Through an AI Literacy Module. AAAI.