On this page

Prompt engineering — the practice of designing and refining inputs to large language models to achieve desired outputs. In education, prompt engineering serves dual roles: as a learner skill (students must learn to prompt effectively) and as a system design lever (developers craft prompts that shape AI tutoring behavior).

Questions to Consider

  • You've likely typed a prompt into an AI tool recently. Now consider this: the way you phrased it isn't neutral — it may reveal how you planned, thought, and allocated your effort. What might your own prompting habits say about how you approach problems?
  • A study found that users who phrase requests skillfully systematically get better output than those expressing the same intent less adroitly. If you accept that 'prompt privilege' is real, is fair access to AI best fixed by teaching everyone to prompt better, or by redesigning the system to not demand that skill — and what are the trade-offs of each?
  • Is prompting a 'trick' to be memorized, or a genuine intellectual skill? One line of research treats it as professional judgment within a discipline (journalism, law, medicine); another treats it as a core of AI literacy. Which view aligns with your own experience of what actually separates good prompts from bad ones?
  • Well-designed prompts can scaffold student thinking, while poorly used ones can encourage cognitive offloading. Can you recall a moment when an AI answer did the thinking for you? What about the prompt — or your intent — made that happen, and could it have been designed to do the opposite?
  • Prompting is both a learner skill and a system-design lever: some tutors now automatically route and select prompts for the user. As prompting moves from the user to the system, what do students lose — and what do they gain?
  • Set a small goal before you read: after learning about prompt engineering, decide on one concrete way you'll change how you write prompts in your own work, and what result you'll check to know it worked.

Introduction

Prompt engineering is central to effective Generative AI use in education. Unlike traditional programming interfaces, LLMs respond to natural language — but the quality, accuracy, and pedagogical value of those responses depend heavily on prompt design. Research in this knowledge base reveals that prompting is not a neutral act: it reflects how students think, plan, and allocate cognitive effort. Miles, Haber-Curran and Arar (2026) sharpen what the term covers by distinguishing prompt engineering, the technical optimization of inputs for performance, from prompt literacy, the rhetorical, ethical and reflective work of clarifying purpose, reading output critically and revising with stated reasons. Their Prompt Literacy Cycle (Clarify Purpose, Craft the Prompt, Engage with Output, Refine the Prompt, Reflect) and a sample process rubric make the distinction teachable, and they argue that instruction which optimizes outputs alone leaves the ethical and epistemological dimensions of LLM use untouched.

How prompt engineering appears in the research

  • Prompting as cognitive trace: Misiejuk et al. show that prompt patterns reveal cognitive offloading — high-quality work uses context-rich, polite, and instructional prompts; low-quality work shows reactive disagreement without domain grounding

  • Prompting as literacy: Tracing GenAI literacy and K-12 prompting literacy research frame prompting as a core AI Literacy component

  • Prompting as system design: CoTAL uses human-in-the-loop prompt engineering for formative assessment scoring; anchor-based prompting improves automated essay scoring

  • Adaptive prompt routing: Learning to Prompt treats prompt selection as part of the tutoring system itself — subject-aware prompt routing over 14 pedagogical features, where a stochastic router selects the best prompt per conversation. This shifts prompting from a learner skill into an adaptive system-design lever, improving engagement and efficiency (28.1% vs 19.6% exercise conversion in a real-world A/B test).

  • Prompt modalities: Voice vs. text input research examines whether prompting modality affects learning outcomes

  • Scaffolded prompting: Guided LLM scaffolding and critical engagement scaffolding teach structured prompting as a learning intervention

  • Prompt privilege and equity: Jin et al. show prompting expertise is unevenly distributed — users who phrase requests skillfully systematically get better output than those expressing the same intent less adroitly. Their Prompt Equity Transformer shifts prompt optimization from the user to the AI system, arguing that equitable output should be engineered into the model rather than demanded of novices.

  • Prompting as situated professional judgment. Beyond literacy and system design, prompting can be framed as a disciplinary practice. The Dierickx et al. taxonomy for journalism treats task definition and prompting as a form of professional judgment exercised within a domain's epistemic and ethical norms — translating journalistic work into explicit tasks (newsgathering → sensemaking → editing → publication/distribution) makes assumptions, priorities, and ethical considerations visible, and turns prompting into a pedagogical tool for critical AI literacy. Its logic transfers to other knowledge-intensive professions (law, medicine, public policy).

  • Prompt design as instructional specification. Neto and colleagues (2026) find in their systematic review of GenAI in healthcare education that prompt design functions as a form of instructional specification, encoding the cognitive targets and quality criteria implicit in expert authoring — yet only 34.8% of studies aligned generated content with instructional frameworks and only 34.8% reported prompting in enough detail to reproduce. Looi, Liu, and Sun (2026) further show how prompt architecture can embed pedagogical rules (correctness gates, anti-spoiler boundaries, goodbye gates) to constrain Large Language Models (LLMs) tutoring behavior in procedural domains.

  • Rubric-guided and role-aware prompting. Yaşar et al. (2026) showed that rubric-guided prompting — treating the rubric as a semantic interface between human pedagogical intent and machine inference — drove LLM–human agreement on student design work from 54.75% to 81.25% (Cronbach's Alpha 0.393 → 0.798). Rubrics engineered for LLMs must balance precision and flexibility: too vague invites free interpretation, too rigid reduces the model to pattern-matching. Role-aware prompting — evaluating the same artifact under instructor, peer-reviewer, and grant-reviewer prompts — produced qualitatively distinct, epistemically different feedback, showing that prompt design shapes not just accuracy but the evaluative stance of the output.

  • Context-aware prompting for assessment. Context-aware prompting of pre-trained language models automates the coding of collaborative problem-solving skills from process data, modeling dependencies between behavior codes and fusing cognitive and social abilities. This enables structured CPS analysis at scale and in real time, overcoming the labor intensity of manual coding schemes.

  • Role-based templates and quality rubrics for teacher planning. Luo and Tahir (2025) empirically develop a prompt framework for children's STEAM arts lesson planning that pairs a Role (R) – Instructions (I) – End Goal (E) template (adapted from RISEN) with a "four points and one line" optimization rubric — standardized, practical, engaging, and complete, plus an extension dimension. Applying the rubric to critique and refine prompts kept generated plans acceptable to practicing art teachers (mean ratings above 4/5) while exposing recurring gaps (personalization, child-safety constraints, cultural bias) that plain one-shot prompting left unaddressed — showing prompt templates plus explicit evaluation criteria function as a quality-control scaffold for classroom generation.

  • Role and constraint design as the independent variable. Wang et al. (2026) compare two agents built on the same model and platform at temperature 0.3 whose only difference is how the prompt specifies role, skills, and constraints: an expert teacher agent answering from a bounded textbook knowledge source versus an empathic student-centered agent scripted to diagnose Misconceptions about AI and check comprehension. The role difference alone shifted learning performance, cognitive load, flow experience, and perceived empathy, showing that role specification is an instructional-design decision with measurable effects rather than a stylistic flourish (Pedagogical Agent).

Connections to broader concepts

Prompt engineering connects to Scaffolding — well-designed prompts can scaffold student thinking rather than bypass it. It intersects with Metacognition and AI Literacy, as effective prompting requires understanding both the AI's capabilities and one's own learning goals. The Cognitive Offloading research directly links prompt quality to whether AI use supports or undermines learning.

  • Writing skill drives prompting, and both predict Vibe Coding success. In a preregistered CHI 2026 study (N=100), Thorgeirsson, Weidmann & Su found that written-communication proficiency predicted GUI-oriented vibe-coding performance (r = .29), with human-graded prompt quality mediating the link — response-process evidence that clear, structured prose translates into better natural-language programming prompts. Both writing skill and CS achievement were independent predictors, and CS achievement (r = .39) carried roughly twice the unique variance, so improving prompting alone is unlikely to fully substitute for programming fundamentals in LLM-native development.
  • Prompting strategy predicts performance. An empirical study of 128 engineering students found that AI Query Efficiency (clear, well-structured prompts) and AI-Driven Problem Solving (strategic integration of AI output into reasoning) were the strongest predictors of academic success — even after controlling for GPA — indicating prompting is a teachable skill that shapes how effectively students learn with AI.
  • A usable taxonomy, and which prompt categories actually pay off. Jacobsen et al. (2026) translate technical strategies into the 3K model (Kontext, Kernauftrag, Klarheit — context, core task, clarity): eleven practice-oriented categories, each with a good/average/suboptimal rubric, and each tested as an experimental variation on feedback generated for pre-service teachers' learning goals. Domain-specific technical language was the decisive category — replacing subject terminology with everyday paraphrases significantly reduced feedback quality across three models (β = −0.412) — while adding concrete examples and removing the chain-of-thought instruction produced no significant difference from the baseline in the first study; examples did help once the analysis was rerun with the best-performing model-prompt combinations (β = 0.52). Prompt quality and model choice together explained 42.8% of the variance in rated feedback quality, which is the paper's case that prompt engineering is a measurable and teachable competency rather than a stylistic preference — and that its categories are not interchangeable in effect size.

Connected Concepts

Connected Articles

Connected Resources

  • EduGems
    'A growing library of pre-made Google Gemini 'Gems' — shareable, pre-written prompts for planning, feedback, forms, images and AI-use reflection.

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.